AI fitness assisting method and system based on neural network

Through the AI ​​fitness assistance system based on neural networks, the deep neural network model is used to identify and monitor fitness movements, the problem of irregular fitness movements is solved, real-time feedback and motion adjustment are achieved, and the fitness effect and safety are improved.

CN119942646APending Publication Date: 2025-05-06QUJING MEDICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510036309.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve automated monitoring and real-time feedback of fitness movements, resulting in irregular fitness movements, inability to achieve the expected fitness effects, and may even cause problems such as muscle strains.

Method used

The AI ​​fitness assistance system based on neural network is adopted to obtain standard movement data and real-time fitness videos of fitness movements through the image acquisition module. The image analysis module extracts the relative position matrix of key human feature points, establishes a deep neural network model, recognizes and detects real-time fitness movements, and generates movement standard detection information and fitness movement prompt information.

Benefits of technology

It realizes accurate identification and monitoring of fitness movements, provides instant feedback, helps fitness workers adjust their movement postures, improve training effects, and avoid injuries. At the same time, it reduces the cost of fitness guidance and improves scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942646A_ABST
    Figure CN119942646A_ABST
Patent Text Reader

Abstract

The invention provides an AI fitness assisting method and system based on a neural network, and relates to the field of image data processing, and the method comprises the steps: extracting a standard feature point relative position matrix corresponding to each standard motion image according to a plurality of key human body feature points, and building a motion recognition model; extracting a real-time feature point relative position matrix corresponding to each real-time fitness image according to the plurality of key human body feature points; through an action recognition model, according to the real-time feature point relative position matrix corresponding to each real-time body-building image, determining a target body-building action from the plurality of body-building actions; according to the standard feature point relative position matrix corresponding to each standard motion image of the target body-building motion and the real-time feature point relative position matrix corresponding to each real-time body-building image, motion standard detection information and body-building motion prompt information are generated. And fitness prompts can be provided for the user in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing, and in particular to an AI fitness assistance method and system based on a neural network. Background Art

[0002] With the improvement of living standards, maintaining a healthy body has become a basic pursuit of people. At present, the fast pace of life and high pressure in cities, especially for office workers, have led to a lack of exercise for more and more people, and the pressure cannot be released, resulting in a sub-healthy state of the body. As people pay more attention to health, running can no longer meet people's higher fitness needs. People hope to exercise with professional equipment and a comfortable fitness environment, and can plan and guide fitness movements through fitness coaches. However, due to the limited number of fitness coaches, fitness coaches cannot supervise the movements of each student. At the same time, due to the high cost of fitness coaches in gyms, fitness personnel usually exercise on their own, which is prone to irregular fitness movements, which in turn leads to the failure to achieve the expected fitness effect, and even causes problems such as muscle strains among fitness personnel.

[0003] Therefore, it is necessary to provide an AI fitness assistance method and system based on a neural network to realize automatic monitoring of fitness movements and provide users with fitness tips in a timely manner. Summary of the invention

[0004] The present invention provides an AI fitness assistance system based on a neural network, comprising: an image acquisition module, used to acquire standard action data of a plurality of fitness actions, wherein the standard action data of the fitness actions include standard action images of a plurality of action nodes in a fitness cycle; an image analysis module, used to extract a relative position matrix of standard feature points corresponding to each of the standard action images according to a plurality of key human feature points, and to establish an action recognition model according to the relative position matrix of standard feature points corresponding to each of the standard action images, wherein the relative position matrix of standard feature points includes the relative positions of any two adjacent key human feature points extracted from the standard action images, and the action recognition model is a deep neural network model; the image acquisition module is also used to acquire a real-time fitness video, and extract a plurality of real-time fitness images from the real-time fitness video ; The image analysis module is also used to extract the real-time feature point relative position matrix corresponding to each of the real-time fitness images according to the multiple key human feature points, wherein the real-time feature point relative position matrix includes the relative positions of any two adjacent key human feature points extracted from the standard action image; the action detection module is used to determine the target fitness action from the multiple fitness actions according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model; the action detection module is also used to generate action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each of the real-time fitness images; the action prompt module is used to generate fitness action prompt information according to the action standard detection information.

[0005] Furthermore, the image analysis module establishes an action recognition model according to a relative position matrix of standard feature points corresponding to each of the standard action images, including: grouping the multiple fitness actions according to the relative position matrix of standard feature points corresponding to multiple standard action images of each of the fitness actions to determine multiple action groups; for each of the action groups, obtaining training samples of the multiple fitness actions included in the action group, and establishing and training the action recognition model corresponding to the action group based on the training samples of the multiple fitness actions included in the action group.

[0006] Further, the image analysis module groups the multiple fitness movements according to the relative position matrix of standard feature points corresponding to the multiple standard action images of each of the fitness movements, and determines multiple action groups, including: for each fitness movement, determining multiple target key human feature points corresponding to the fitness movement according to the relative position matrix of standard feature points corresponding to the multiple standard action images of the fitness movement; classifying the multiple fitness movements according to the multiple target key human feature points corresponding to each fitness movement, and determining multiple action classes; for each action class, grouping the multiple fitness movements included in the action class according to the relative position matrix of standard feature points corresponding to the multiple standard action images of each fitness movement included in the action class and the corresponding multiple target key human feature points, and determining the action group included in the action class.

[0007] Further, the image analysis module determines multiple target key human feature points corresponding to the fitness action based on a relative position matrix of standard feature points corresponding to multiple standard action images of the fitness action, including: for each key human feature point, determining the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image based on the relative position matrix of standard feature points corresponding to multiple standard action images of the fitness action, calculating a first posture change parameter corresponding to the key human feature point based on the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image, and judging whether the key human feature point is the target key human feature point corresponding to the fitness action based on the first posture change parameter corresponding to the key human feature point.

[0008] Furthermore, the image analysis module classifies the multiple fitness movements according to the multiple target key human feature points corresponding to each fitness movement, and determines multiple action classes, including: for any two fitness movements, calculating the overlap of the target key human feature points of the two fitness movements according to the multiple target key human feature points corresponding to the two fitness movements; classifying the multiple fitness movements according to the overlap of the target key human feature points of the two fitness movements through a clustering algorithm, and determining multiple action classes; for each action class, determining the multiple class key human feature points corresponding to the action class according to the multiple target key human feature points corresponding to each fitness movement included in the action class.

[0009] Further, the image analysis module groups the multiple fitness movements included in the action class according to the relative position matrix of standard feature points corresponding to multiple standard action images of each fitness movement included in the action class and the corresponding multiple target key human feature points, and determines the action group included in the action class, including: for any two fitness movements included in the action class, according to the first posture change parameter of each class key human feature point corresponding to each fitness movement, calculating the action similarity of the two fitness movements; through a clustering algorithm, according to the action similarity of any two fitness movements included in the action class, grouping the multiple fitness movements included in the action class, and determining the action group included in the action class; for each action group, according to the first posture change parameter of each class key human feature point corresponding to each fitness movement included in the action group, determining the group center posture change parameter of each class key human feature point corresponding to the action group.

[0010] Further, the action detection module determines the target fitness action from the multiple fitness actions according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model, including: determining the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images, calculating the second posture change parameter corresponding to the key human feature point according to the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image, and calculating the second posture change parameter corresponding to the key human feature point according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images. The second posture change parameter of the key human feature point is used to determine whether the key human feature point is a real-time key human feature point; according to multiple real-time key human feature points and multiple class key human feature points corresponding to each action class, a target action class is determined from the multiple action classes; according to the second posture change parameter corresponding to each real-time key human feature point and the group center posture change parameter of each class key human feature point corresponding to each group included in the target action class, a target action group is determined from the multiple action groups included in the target action class; through the action recognition model corresponding to the target action group, a target fitness action is determined from the multiple fitness actions according to the second posture change parameter corresponding to each real-time key human feature point.

[0011] Further, the action detection module generates action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image, including: generating a plurality of image groups to be compared through an image group generation model according to the real-time feature point relative position matrix and time tag corresponding to each real-time fitness image and the standard feature point relative position matrix and time tag corresponding to each standard action image of the target fitness action, wherein the image groups to be compared include a real-time fitness image and a standard action image, and the number of the image groups to be compared is consistent with the number of the standard action images; for each image group to be compared, calculating the action difference parameter corresponding to the image group to be compared according to the real-time feature point relative position matrix corresponding to the real-time fitness image included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action image, wherein the action standard detection information includes the action difference parameter corresponding to each image group to be compared.

[0012] Further, the action detection module calculates the action difference parameters corresponding to the image group to be compared according to the real-time feature point relative position matrix corresponding to the real-time fitness images included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action images, including: for each key human feature point, according to the real-time feature point relative position matrix corresponding to the real-time fitness images included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action images, calculating the single point difference parameter corresponding to the key human feature point; according to the single point difference parameter corresponding to each key human feature point, calculating the action difference parameter corresponding to the image group to be compared.

[0013] The present invention provides an AI fitness assistance method based on a neural network, which is applied to the above-mentioned AI fitness assistance system based on a neural network, comprising: obtaining standard action data of multiple fitness actions, wherein the standard action data of the fitness actions include standard action images of multiple action nodes in a fitness cycle; extracting a standard feature point relative position matrix corresponding to each of the standard action images according to multiple key human feature points; establishing an action recognition model according to the standard feature point relative position matrix corresponding to each of the standard action images, wherein the standard feature point relative position matrix includes the relative positions of any two adjacent key human feature points extracted from the standard action images, and the action recognition model is a deep neural network model; obtaining real-time fitness video, and extracting the standard feature point relative position matrix from the standard action image; and obtaining a standard fitness video. Extract multiple real-time fitness images from the real-time fitness video; extract a real-time feature point relative position matrix corresponding to each of the real-time fitness images according to the multiple key human feature points, wherein the real-time feature point relative position matrix includes the relative positions of any two adjacent key human feature points extracted from the standard action image; determine a target fitness action from the multiple fitness actions according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model; generate action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each of the real-time fitness images; generate fitness action prompt information according to the action standard detection information.

[0014] Compared with the prior art, the neural network-based AI fitness assistance method and system provided by the present invention has at least the following beneficial effects:

[0015] 1. By acquiring standard action data and real-time fitness videos, the relative position matrix of key human feature points can be extracted, and then an action recognition model can be established. With this model, real-time fitness actions can be accurately identified and determined, and compared with standard actions to ensure the accuracy of the fitness practitioner's actions. This helps to avoid injuries that may be caused by incorrect fitness postures and improve training effects. It can provide instant feedback for each fitness practitioner's real-time actions. According to the action standard detection information, personalized fitness action prompt information is generated to help fitness practitioners adjust their actions in time to achieve the best training state. This personalized guidance is more efficient and convenient than traditional fitness coach guidance. With the help of AI technology, fitness actions can be analyzed and feedback can be provided in real time without waiting for the coach's guidance or manually checking the actions. This not only saves time, but also improves the efficiency and fun of fitness. At the same time, automation and intelligence also enhance the interactivity and participation of the fitness process. By collecting and analyzing a large amount of fitness action data, the action recognition model and training guidance strategy can be continuously optimized. Compared with traditional fitness coaches, it has lower costs and higher scalability. This makes it easier for high-quality fitness guidance services to be popularized to a wider range of fitness people and promote the dissemination and development of fitness culture.

[0016] 2. By extracting the relative position matrix of key human feature points in standard action images, fitness actions are grouped and classified based on these feature points. This feature point-based method can capture subtle changes in human actions, thereby improving the accuracy of action recognition. An action recognition model is established and trained for each action group, so that each model can be optimized for a specific type of action. This group training method helps to improve recognition efficiency and reduce the misrecognition rate.

[0017] 3. Through the image group generation model, the real-time fitness images are paired with the standard action images to form a group of images to be compared. This method ensures that each frame of real-time fitness action can be accurately compared with the corresponding standard action, thereby improving the accuracy of action detection. For each group of images to be compared, the single-point difference parameters of key human feature points are calculated, and the action difference parameters are derived accordingly. This detailed difference analysis helps the system to more accurately identify the movement deviations of the fitness person and provide strong support for subsequent fitness guidance. It can capture the fitness person's movements in real time, compare them with standard movements, and generate movement standard detection information. This real-time feedback mechanism allows fitness people to immediately understand whether their movements are standard, so that they can adjust their movements and postures in time to improve training effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0019] Figure 1 is a module schematic diagram of an AI fitness assistance system based on a neural network according to some embodiments of this specification;

[0020] Figure 2 It is a flowchart of an AI fitness assistance method based on a neural network as shown in some embodiments of this specification. DETAILED DESCRIPTION

[0021] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of this specification. For ordinary technicians in this field, this specification can also be applied to other similar scenarios based on these drawings without creative work. Unless it is obvious from the language environment or otherwise explained, the same reference numerals in the figures represent the same structure or operation.

[0022] Figure 1 is a module schematic diagram of an AI fitness assistance system based on a neural network according to some embodiments of this specification, such as Figure 1 As shown, a neural network-based AI fitness assistance system may include an image acquisition module, an image analysis module, an action detection module and an action prompt module.

[0023] The image acquisition module can be used to acquire standard motion data of a plurality of fitness movements, wherein the standard motion data of the fitness movements include standard motion images of a plurality of motion nodes in a fitness cycle.

[0024] Taking squatting as an example, a squat can be broken down into the following action nodes:

[0025] 1. Preparation stage (time point: 0 seconds - 1 second)

[0026] Action description: Stand with your feet shoulder-width apart or slightly wider, toes slightly outward, keep your body balanced. Put your hands on your waist.

[0027] 2. Inhale and start squatting (time point: 1 second - 3 seconds)

[0028] Action description: Take a deep breath and prepare to squat. At the same time, slowly bend your knees so that your thighs form a certain angle with the ground.

[0029] 3. Squat to the lowest point (time point: 3 seconds - 5 seconds)

[0030] Action description: Continue squatting until your thighs are parallel to or slightly lower than the ground, and your hips should move backward.

[0031] 4. Exhale and prepare to stand up (time point: 5 seconds to 6 seconds)

[0032] Action description: After squatting to the lowest point, start exhaling to prepare for standing up.

[0033] 5. Stand up and stand again (time point: 6 seconds to 8 seconds)

[0034] Action description: Using the strength of your thigh and buttocks muscles, slowly stand up and return to the starting standing position.

[0035] Each action node collects a standard action image.

[0036] The image analysis module can be used to extract the relative position matrix of standard feature points corresponding to each standard action image based on multiple key human feature points, and establish an action recognition model based on the relative position matrix of standard feature points corresponding to each standard action image, wherein the relative position matrix of standard feature points includes the relative positions of any two adjacent key human feature points extracted from the standard action image, and the action recognition model is a deep neural network model.

[0037] By way of example only, the plurality of key human feature points may include toes, ankles, knees, waist, wrists, elbows, shoulders, eyes, ears, and the like.

[0038] A row of the standard feature point relative position matrix corresponds to a key human feature point, and the row vector of the standard feature point relative position matrix is ​​composed of the relative position of the corresponding key human feature point and its adjacent key human feature point. The relative position may include the straight-line distance between the two key human feature points and the angle between the straight line connecting the two key human feature points and the horizontal line.

[0039] As an example only, the key human feature points adjacent to the waist include the knees and shoulders, so the row vector of the waist corresponding to the standard feature point relative position matrix is ​​(D (waist,knee) , A (waist,knee) , D (shoulder,waist) , A (shoulder,waist) ), where D (waist,knee) is the straight-line distance between waist and knee, A (waist,knee) D is the angle between the straight line between the waist and the knee and the horizontal line. (shoulder,waist) is the straight-line distance between waist and shoulder, A (shoulder,waist) It is the angle between the straight line between waist and shoulder and the horizontal line.

[0040] The adjacent relationship between multiple key human feature points can be manually defined in advance.

[0041] Gaussian filtering and moving average method are used to perform image smoothing and data smoothing, as well as de-jittering processing, to improve the stability and accuracy of recognition, satisfying the following formula:

[0042]

[0043] Where G(x, y) is the Gaussian kernel function and σ is the standard deviation.

[0044] Graph Convolution Neural Networks (GCNNs) are used to detect and connect key human feature points in standard action images, satisfying the following formula:

[0045] (I*K)(x,y)=Em En I(m,n)*K(xm,yn)

[0046] Where, I: input image or feature map. It is a two-dimensional matrix in which each element I(m, n) represents the pixel value or feature value of the image at position (m, n). K: convolution kernel (or filter). It is also a two-dimensional matrix used to slide on the image and perform element-level multiplication and summation operations to extract features. K(xm, yn) represents the value of the convolution kernel at position (xm, yn). (x, y): coordinates of the output feature map. This is a certain position of the new feature map generated after the convolution operation. (m, n): convolution kernel coordinates (offset relative to the current output position (x, y)). These coordinates are used to traverse all elements of the convolution kernel and multiply them with the corresponding positions on the input image or feature map. *: The symbol here indicates the convolution operation, that is, the corresponding elements of the input image and the convolution kernel are multiplied first, and then the results are summed.

[0047] In some embodiments, the image analysis module establishes an action recognition model according to the relative position matrix of standard feature points corresponding to each standard action image, including:

[0048] Grouping multiple fitness movements according to a relative position matrix of standard feature points corresponding to multiple standard movement images of each fitness movement to determine multiple movement groups;

[0049] For each action group, training samples of multiple fitness actions included in the action group are obtained, and based on the training samples of the multiple fitness actions included in the action group, an action recognition model corresponding to the action group is established and trained.

[0050] In some embodiments, the image analysis module groups multiple fitness movements according to the relative position matrix of standard feature points corresponding to multiple standard movement images of each fitness movement to determine multiple movement groups, including:

[0051] For each fitness action, determining multiple target key human body feature points corresponding to the fitness action according to a relative position matrix of standard feature points corresponding to multiple standard action images of the fitness action;

[0052] Classifying multiple fitness movements according to multiple target key human feature points corresponding to each fitness movement, and determining multiple movement classes;

[0053] For each action class, the multiple fitness actions included in the action class are grouped according to the relative position matrix of standard feature points corresponding to multiple standard action images of each fitness action included in the action class and the corresponding multiple target key human feature points to determine the action group included in the action class.

[0054] In some embodiments, the image analysis module determines a plurality of target key human feature points corresponding to the fitness action according to a relative position matrix of standard feature points corresponding to a plurality of standard action images of the fitness action, including:

[0055] For each key human feature point, the relative position matrix of the standard feature points corresponding to multiple standard action images of the fitness action is used to determine the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image; the first posture change parameter corresponding to the key human feature point is calculated based on the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image; and based on the first posture change parameter corresponding to the key human feature point, it is determined whether the key human feature point is the target key human feature point corresponding to the fitness action.

[0056] Specifically, when the number of adjacent key human feature points of a key human feature point is 1, the first posture change parameter corresponding to the key human feature point can be calculated according to the following formula:

[0057]

[0058] Among them, V i is the first posture change parameter corresponding to the i-th key human feature point, D (i,adjacent,n) is the straight-line distance between the ith key human feature point and the adjacent key human feature point of the ith key human feature point corresponding to the nth standard image, A (i,adjacent) is the angle between the straight line connecting the ith key human feature point and the nth standard image corresponding to the ith key human feature point and the horizontal line, and N is the total number of standard action images of the fitness action.

[0059] When the number of adjacent key human feature points of a key human feature point is 2, the first posture change parameter corresponding to the key human feature point can be calculated according to the following formula:

[0060]

[0061] Among them, V (i,m) is the posture change parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point, D (i,m,n) is the linear distance change parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point, A (i,m,n) is the angle variation parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point.

[0062] The key human feature points whose first posture change parameters are greater than the first posture change parameter threshold may be used as target key human feature points corresponding to the fitness action.

[0063] In some embodiments, the image analysis module classifies multiple fitness movements according to multiple target key human feature points corresponding to each fitness movement, and determines multiple movement categories, including:

[0064] For any two fitness movements, the overlap degree of the target key human feature points of the two fitness movements is calculated according to the multiple target key human feature points corresponding to the two fitness movements;

[0065] Classifying the multiple fitness movements according to the overlap of target key human feature points of two fitness movements through a clustering algorithm (e.g., a K-means clustering algorithm, etc.) to determine multiple movement classes;

[0066] For each action class, a plurality of class key human feature points corresponding to the action class are determined according to a plurality of target key human feature points corresponding to each fitness action included in the action class.

[0067] For example, the overlap of target key human feature points of two fitness actions can be calculated according to the following formula:

[0068]

[0069] Among them, R (i,j) is the overlap degree of the target key human feature points between the i-th fitness action and the j-th fitness action, N i∩j is the total number of target key human feature points included in the intersection of the target key human feature points of the i-th fitness action and the target key human feature points of the j-th fitness action, N i is the total number of target key human feature points for the i-th fitness action, N j is the total number of target key human feature points of the jth fitness action, max(N i , N j ) is to take N i and Nj The larger value in .

[0070] For each action class, a plurality of target key human feature points corresponding to each fitness action included in the action class may be unioned to obtain a plurality of candidate key human feature points corresponding to the action class. For each candidate key human feature point, the number of fitness actions in the action class that use the candidate key human feature point as a target key human feature point may be calculated, and when the number is greater than a number threshold, the candidate key human feature point may be used as a class key human feature point corresponding to the action class.

[0071] In some embodiments, the image analysis module groups multiple fitness actions included in the action class according to the standard feature point relative position matrix corresponding to multiple standard action images of each fitness action included in the action class and the corresponding multiple target key human feature points, and determines the action group included in the action class, including:

[0072] For any two fitness actions included in the action class, calculating the action similarity of the two fitness actions according to the first posture change parameter of each class key human feature point corresponding to each fitness action;

[0073] By using a clustering algorithm (e.g., a K-means clustering algorithm, etc.), according to the action similarity between any two fitness actions included in the action class, a plurality of fitness actions included in the action class are grouped to determine an action group included in the action class;

[0074] For each action group, the group center posture change parameter corresponding to each class key human feature point of the action group is determined according to the first posture change parameter of each class key human feature point corresponding to each fitness action included in the action group.

[0075] Specifically, the action similarity of two fitness actions can be calculated according to the following formula:

[0076]

[0077] Among them, S (i,j) is the action similarity between the ith fitness action and the jth fitness action included in the action class, P1 is a preset parameter, P1 is greater than 0, V (i,e) is the first posture change parameter of the e-th class key human feature point corresponding to the i-th fitness action included in the action class, V (j,e) is the first posture change parameter of the e-th class key human feature point corresponding to the j-th fitness action included in the action class, and E is the total number of class key human feature points corresponding to the action class.

[0078] For each class of key human feature points, the first posture change parameters of the key human feature points of this class corresponding to each fitness action included in the action group can be averaged as the group center posture change parameters of the action group corresponding to the key human feature points of this class.

[0079] The action recognition model can be a lightweight deep neural network model (MobileNetV2), which enables a neural network-based AI fitness assistance system to run smoothly on resource-constrained devices (such as smartphones and tablets), while improving computing efficiency and response speed. In addition, the MobileNetV2 model is optimized using pruning technology to further improve the inference speed and low resource usage.

[0080] The expression of the loss function is minimized by iteratively updating the model's parameters:

[0081] Loss function:

[0082]

[0083] L(θ) is the loss function, θ is the model parameter, N is the number of samples, y i is the actual value, is the model prediction value, is the loss of a single sample.

[0084] The L1 norm expression used for weight pruning is:

[0085]

[0086] Among them, suppose the weight matrix of a certain layer is W = [w1, w2, ...., w n ],w ij represents the jth weight of the i-th weight matrix.

[0087] The image acquisition module is also used to acquire real-time fitness videos and extract multiple real-time fitness images from the real-time fitness videos.

[0088] Specifically, the image acquisition module may include an image acquisition device for acquiring real-time fitness videos.

[0089] The image analysis module is also used to extract a real-time feature point relative position matrix corresponding to each real-time fitness image based on multiple key human feature points, wherein the real-time feature point relative position matrix includes the relative positions of any two adjacent key human feature points extracted from the standard action image.

[0090] The action detection module can be used to determine a target fitness action from multiple fitness actions according to a real-time feature point relative position matrix corresponding to each real-time fitness image through an action recognition model.

[0091] In some embodiments, the action detection module determines a target fitness action from a plurality of fitness actions according to a real-time feature point relative position matrix corresponding to each real-time fitness image through an action recognition model, including:

[0092] Determine the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image according to the relative position matrix of the real-time feature points corresponding to each real-time fitness image, calculate the second posture change parameter corresponding to the key human feature point according to the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image, and judge whether the key human feature point is a real-time key human feature point according to the second posture change parameter corresponding to the key human feature point;

[0093] Determine a target action class from the multiple action classes according to the multiple real-time key human feature points and the multiple class key human feature points corresponding to each action class;

[0094] Determine a target action group from a plurality of action groups included in the target action class according to a second posture change parameter corresponding to each real-time key human feature point and a group center posture change parameter of each group included in the target action class corresponding to each class key human feature point;

[0095] A target fitness action is determined from a plurality of fitness actions through an action recognition model corresponding to the target action group and according to a second posture change parameter corresponding to each real-time key human feature point.

[0096] Specifically, the method of calculating the second posture change parameter corresponding to the key human feature point is similar to the method of calculating the first posture change parameter corresponding to the key human feature point, which will not be repeated here. The key human feature point whose corresponding second posture change parameter is greater than the second posture change parameter threshold can be used as the real-time key human feature point.

[0097] For each action class, the overlap of key human feature points can be calculated based on the real-time key human feature points and multiple class key human feature points corresponding to the action class, and the action class whose overlap of key human feature points is greater than the key human feature point overlap threshold is taken as the target action class.

[0098] For each action group included in the target action class, the matching degree corresponding to the action group can be calculated based on the second posture change parameter corresponding to each real-time key human feature point and the group center posture change parameter of each group included in the target action class corresponding to each class key human feature point, and the action group with a matching degree greater than the matching degree threshold is used as the target action group.

[0099] For example, the matching degree corresponding to the action group can be calculated according to the following formula:

[0100]

[0101] Among them, M i is the matching degree corresponding to the i-th action group, P2 is the preset parameter, P2 is greater than 0, V (current,e) is the second posture change parameter corresponding to the key human feature point of the e-th class, V (i,e) is the group center posture change parameter of the e-th class of key human feature points corresponding to the ith action group.

[0102] Through the action recognition model corresponding to the target action group, according to the second posture change parameter corresponding to each real-time key human feature point, the probability of each fitness action included in the target action group being the target fitness action is determined, and the fitness action with the highest probability is used as the target fitness action.

[0103] The action detection module is also used to generate action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image.

[0104] In some embodiments, the action detection module generates action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image, including:

[0105] Generate multiple groups of image groups to be compared by using an image group generation model, according to a real-time feature point relative position matrix and a time label corresponding to each real-time fitness image and a standard feature point relative position matrix and a time label corresponding to each standard action image of a target fitness action, wherein the image group to be compared includes a real-time fitness image and a standard action image, the number of image groups to be compared is consistent with the number of standard action images, the image group generation model may be a convolutional neural network model, and in the image group to be compared, the action nodes corresponding to the real-time fitness images identified by the image group generation model are consistent with the action nodes corresponding to the standard action images;

[0106] For each image group to be compared, the action difference parameters corresponding to the image group to be compared are calculated based on the real-time feature point relative position matrix corresponding to the real-time fitness images included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action images, wherein the action standard detection information includes the action difference parameters corresponding to each image group to be compared.

[0107] In some embodiments, the action detection module calculates the action difference parameter corresponding to the image group to be compared according to the real-time feature point relative position matrix corresponding to the real-time fitness image included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action image, including:

[0108] For each key human feature point, a single point difference parameter corresponding to the key human feature point is calculated according to a real-time feature point relative position matrix corresponding to the real-time fitness image and a standard feature point relative position matrix corresponding to the standard action image included in the image group to be compared;

[0109] According to the single-point difference parameters corresponding to each key human feature point, the action difference parameters corresponding to the image group to be compared are calculated.

[0110] Specifically, for each class key human feature point corresponding to the target action class, the real-time relative position of the class key human feature point and each adjacent key human feature point of the class key human feature point can be determined according to the real-time fitness image included in the image group to be compared, and the standard relative position of the class key human feature point and each adjacent key human feature point of the class key human feature point can be determined according to the standard action image included in the image group to be compared; based on the real-time relative position of the class key human feature point and each adjacent key human feature point of the class key human feature point and the standard relative position of the class key human feature point and each adjacent key human feature point of the class key human feature point, the single-point difference parameter corresponding to the class key human feature point is calculated.

[0111] When the number of adjacent key human feature points of a class key human feature point is 1, the single point difference parameter corresponding to the class key human feature point can be calculated according to the following formula:

[0112]

[0113]

[0114] Among them, D i is the single point difference parameter corresponding to the key human feature point of the i-th class, D (i,m) is the difference parameter between the key human feature point of the ith class and the mth adjacent key human feature point of the ith class in the image group to be compared, D ((i,m),1) is the straight-line distance between the key human feature point of the ith class and the mth adjacent key human feature point of the ith class in the real-time fitness image to be compared, D ((i,m),2) A is the straight-line distance between the i-th class key human feature point and the m-th adjacent key human feature point of the i-th class key human feature point in the standard action image included in the image group to be compared, ((i,m),1) is the angle between the horizontal line and the straight line connecting the i-th class key human feature point and the m-th adjacent key human feature point of the i-th class key human feature point in the real-time fitness image to be compared, A ((i,m),2)is the angle between the straight line connecting the i-th class key human feature point and the m-th adjacent key human feature point of the i-th class key human feature point in the standard fitness image included in the compared image group and the horizontal line, P3 is a preset parameter, P3 is greater than 0, and M is the total number of adjacent key human feature points of the i-th class key human feature point.

[0115] When the number of adjacent key human feature points of a key human feature point is 2, the first posture change parameter corresponding to the key human feature point can be calculated according to the following formula:

[0116]

[0117] Among them, V (i,m) is the posture change parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point, D (i,m) is the linear distance change parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point, A (i,m) is the angle variation parameter between the ith key human feature point and the mth adjacent key human feature point of the ith key human feature point.

[0118] The single-point difference parameters corresponding to each class of key human feature points may be weighted summed to calculate the action difference parameters corresponding to the image group to be compared.

[0119] The weight corresponding to each class of key human feature points can be determined by the following process:

[0120] S11, determining the body shape information of the current user according to the multiple real-time fitness images, for example, the size of each part (for example, thigh, waist, calf, arm, shoulder, etc.);

[0121] S12, determining similar users based on the current user's body shape information, for example, representing the current user's body shape information as a vector (or matrix), and then calculating the Euclidean distance, Manhattan distance, cosine similarity, etc. of the vector corresponding to the body shape information of each user in the database to calculate a similarity score, and taking users whose similarity scores are greater than a similarity score threshold as similar users;

[0122] S13. Determine the weights corresponding to the key human feature points of each class through a weight determination model according to historical action standard detection information of similar users, wherein the weight determination model may be a convolutional neural network model.

[0123] The action prompt module can be used to generate fitness action prompt information according to action standard detection information.

[0124] Specifically, when there is at least one action difference parameter corresponding to the to-be-compared image group that is greater than the action difference parameter threshold, the action prompt module can be used to generate fitness action prompt information according to the action standard detection information.

[0125] As an example only, the fitness action prompt information may be generated by the action prompt model according to the action difference parameters corresponding to each to-be-compared image group.

[0126] Example 1: Tips for improper squat movements

[0127] Image group to be compared: The real-time fitness image shows that when the fitness person is doing squats, his back cannot be kept straight and his knees are beyond his toes.

[0128] Movement difference parameters: The back is tilted too much, and the relative position of the knees and toes is inappropriate.

[0129] Fitness exercise tips: "Please pay attention to keep your back upright and avoid leaning forward too much. At the same time, your knees should not exceed your toes when squatting to protect your knee joints."

[0130] Example 2: Tips for Push-ups Not Performing Well

[0131] Image group to be compared: The real-time fitness image shows that when a fitness person is doing push-ups, his elbows are excessively extended outward and his hips are raised.

[0132] Movement difference parameters: elbow angle is too large, and the distance between the hips and the ground is too large.

[0133] Fitness exercise tips: "Please make sure your elbows are close to your sides and avoid over-extending. At the same time, keep your hips close to the ground and do not lift them to ensure standard movements."

[0134] Example 3: Inaccurate tips for yoga cat-cow pose

[0135] Image group to be compared: Real-time fitness images show that when the fitness person is doing the yoga cat-cow pose, the back is not arched enough and the head fails to form a natural curve with the back.

[0136] Movement difference parameters: insufficient back arch and inappropriate relative position of head and back.

[0137] Fitness action prompt: "Please increase the degree of arching of your back and let your head droop naturally to form a smooth curve with your back. At the same time, pay attention to breathing coordination to achieve a better relaxation effect."

[0138] Figure 2 is a flowchart of an AI fitness assistance method based on a neural network according to some embodiments of this specification, such as Figure 2As shown, a neural network-based AI fitness assistance method may include the following steps.

[0139] Step 210, obtaining standard motion data of a plurality of fitness movements, wherein the standard motion data of the fitness movements include standard motion images of a plurality of motion nodes in a fitness cycle;

[0140] Step 220, extracting a relative position matrix of standard feature points corresponding to each standard action image based on the multiple key human feature points;

[0141] Step 230, establishing an action recognition model according to a relative position matrix of standard feature points corresponding to each standard action image, wherein the relative position matrix of standard feature points includes relative positions of any two adjacent key human feature points extracted from the standard action image, and the action recognition model is a deep neural network model;

[0142] Step 240, obtaining a real-time fitness video, and extracting a plurality of real-time fitness images from the real-time fitness video;

[0143] Step 250, extracting a real-time feature point relative position matrix corresponding to each real-time fitness image according to a plurality of key human feature points;

[0144] Step 260, determining a target fitness action from a plurality of fitness actions according to a real-time feature point relative position matrix corresponding to each real-time fitness image through an action recognition model, wherein the real-time feature point relative position matrix includes relative positions of any two adjacent key human feature points extracted from the standard action image;

[0145] Step 270, generating action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image;

[0146] Step 280, generating fitness action prompt information according to the action standard detection information.

[0147] A neural network-based AI fitness assistance method can be applied to the above-mentioned neural network-based AI fitness assistance system. For more description, please refer to the above and will not be repeated here.

[0148] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, as an example and not a limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.

Claims

1. An AI fitness assistance system based on neural network, characterized in that: include: An image acquisition module, used to acquire standard motion data of a plurality of fitness movements, wherein the standard motion data of the fitness movements include standard motion images of a plurality of motion nodes in a fitness cycle; An image analysis module, used to extract a relative position matrix of standard feature points corresponding to each of the standard action images according to a plurality of key human feature points, and to establish an action recognition model according to the relative position matrix of standard feature points corresponding to each of the standard action images, wherein the relative position matrix of standard feature points includes relative positions of any two adjacent key human feature points extracted from the standard action images, and the action recognition model is a deep neural network model; The image acquisition module is also used to acquire a real-time fitness video and extract a plurality of real-time fitness images from the real-time fitness video; The image analysis module is further used to extract a real-time feature point relative position matrix corresponding to each of the real-time fitness images according to the multiple key human feature points, wherein the real-time feature point relative position matrix includes relative positions of any two adjacent key human feature points extracted from the standard action image; an action detection module, configured to determine a target fitness action from the plurality of fitness actions according to a real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model; The action detection module is further used to generate action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image; The action prompt module is used to generate fitness action prompt information according to the action standard detection information.

2. The AI ​​fitness assistance system based on neural network according to claim 1, characterized in that: The image analysis module establishes an action recognition model according to the relative position matrix of standard feature points corresponding to each of the standard action images, including: According to the relative position matrix of the standard feature points corresponding to the plurality of standard action images of each of the fitness actions, the plurality of fitness actions are grouped to determine a plurality of action groups; For each of the action groups, training samples of a plurality of fitness actions included in the action group are obtained, and based on the training samples of the plurality of fitness actions included in the action group, an action recognition model corresponding to the action group is established and trained.

3. The AI ​​fitness assistance system based on neural network according to claim 2, characterized in that: The image analysis module groups the multiple fitness movements according to the relative position matrix of the standard feature points corresponding to the multiple standard movement images of each of the fitness movements to determine multiple movement groups, including: For each fitness action, determining a plurality of target key human body feature points corresponding to the fitness action according to a relative position matrix of standard feature points corresponding to a plurality of standard action images of the fitness action; Classifying the plurality of fitness movements according to the plurality of target key human feature points corresponding to each fitness movement to determine a plurality of movement classes; For each action class, based on the relative position matrix of standard feature points corresponding to multiple standard action images of each fitness action included in the action class and the corresponding multiple target key human feature points, the multiple fitness actions included in the action class are grouped to determine the action group included in the action class.

4. The AI ​​fitness assistance system based on neural network according to claim 3, characterized in that: The image analysis module determines a plurality of target key human body feature points corresponding to the fitness action according to a relative position matrix of standard feature points corresponding to a plurality of standard action images of the fitness action, including: For each key human feature point, the relative position matrix of the standard feature points corresponding to multiple standard action images of the fitness action is used to determine the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image; the first posture change parameter corresponding to the key human feature point is calculated based on the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each standard action image; and based on the first posture change parameter corresponding to the key human feature point, it is determined whether the key human feature point is the target key human feature point corresponding to the fitness action.

5. The AI ​​fitness assistance system based on neural network according to claim 3 or 4, characterized in that: The image analysis module classifies the multiple fitness movements according to the multiple target key human feature points corresponding to each fitness movement, and determines multiple movement categories, including: For any two fitness movements, calculating the overlap degree of the target key human feature points of the two fitness movements according to a plurality of target key human feature points corresponding to the two fitness movements; Classifying the plurality of fitness movements according to the overlap of the target key human feature points of the two fitness movements through a clustering algorithm to determine a plurality of movement classes; For each action class, a plurality of class key human feature points corresponding to the action class are determined according to a plurality of target key human feature points corresponding to each fitness action included in the action class.

6. The neural network-based AI fitness assistance system according to claim 5, characterized in that: The image analysis module groups the multiple fitness movements included in the action class according to the relative position matrix of the standard feature points corresponding to the multiple standard action images of each fitness movement included in the action class and the corresponding multiple target key human feature points, and determines the action group included in the action class, including: For any two fitness actions included in the action class, calculating the action similarity of the two fitness actions according to the first posture change parameter of each class key human feature point corresponding to each fitness action; By using a clustering algorithm, according to the action similarity of any two fitness actions included in the action class, a plurality of fitness actions included in the action class are grouped to determine an action group included in the action class; For each action group, the group center posture change parameter corresponding to each class key human feature point of the action group is determined according to the first posture change parameter of each class key human feature point corresponding to each fitness action included in the action group.

7. The neural network-based AI fitness assistance system according to claim 6, characterized in that: The action detection module determines a target fitness action from the plurality of fitness actions according to a real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model, including: Determine, according to the real-time feature point relative position matrix corresponding to each of the real-time fitness images, the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image; calculate, according to the relative position of the key human feature point and each adjacent key human feature point of the key human feature point corresponding to each real-time fitness image, the second posture change parameter corresponding to the key human feature point; and determine, according to the second posture change parameter corresponding to the key human feature point, whether the key human feature point is a real-time key human feature point; Determine a target action class from the multiple action classes according to the multiple real-time key human feature points and the multiple class key human feature points corresponding to each action class; Determine a target action group from a plurality of action groups included in the target action class according to a second posture change parameter corresponding to each real-time key human feature point and a group center posture change parameter of each group included in the target action class corresponding to each class key human feature point; The target fitness action is determined from the multiple fitness actions through the action recognition model corresponding to the target action group and according to the second posture change parameter corresponding to each real-time key human feature point.

8. The neural network-based AI fitness assistance system according to claim 7, characterized in that: The action detection module generates action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image, including: Generate a plurality of image groups to be compared by using an image group generation model, according to a real-time feature point relative position matrix and a time tag corresponding to each of the real-time fitness images and a standard feature point relative position matrix and a time tag corresponding to each standard action image of the target fitness action, wherein the image groups to be compared include a real-time fitness image and a standard action image, and the number of the image groups to be compared is consistent with the number of the standard action images; For each image group to be compared, the action difference parameters corresponding to the image group to be compared are calculated based on the real-time feature point relative position matrix corresponding to the real-time fitness images included in the image group to be compared and the standard feature point relative position matrix corresponding to the standard action images, wherein the action standard detection information includes the action difference parameters corresponding to each image group to be compared.

9. The neural network-based AI fitness assistance system according to claim 8, characterized in that: The action detection module calculates the action difference parameters corresponding to the image group to be compared according to the real-time feature point relative position matrix corresponding to the real-time fitness image and the standard feature point relative position matrix corresponding to the standard action image included in the image group to be compared, including: For each key human feature point, calculating a single point difference parameter corresponding to the key human feature point according to a real-time feature point relative position matrix corresponding to the real-time fitness image and a standard feature point relative position matrix corresponding to the standard action image included in the image group to be compared; According to the single-point difference parameter corresponding to each key human feature point, the action difference parameter corresponding to the image group to be compared is calculated.

10. An AI fitness assistance method based on neural network, characterized in that: An AI fitness assistance system based on a neural network as described in any one of claims 1 to 9, comprising: Acquire standard motion data of a plurality of fitness movements, wherein the standard motion data of the fitness movements include standard motion images of a plurality of motion nodes in a fitness cycle; Extracting a relative position matrix of standard feature points corresponding to each of the standard action images according to a plurality of key human feature points; Establishing an action recognition model according to a relative position matrix of standard feature points corresponding to each of the standard action images, wherein the relative position matrix of standard feature points includes relative positions of any two adjacent key human feature points extracted from the standard action images, and the action recognition model is a deep neural network model; Acquire a real-time fitness video, and extract a plurality of real-time fitness images from the real-time fitness video; Extracting a real-time feature point relative position matrix corresponding to each of the real-time fitness images according to the multiple key human feature points, wherein the real-time feature point relative position matrix includes relative positions of any two adjacent key human feature points extracted from the standard action images; Determining a target fitness action from the plurality of fitness actions according to a real-time feature point relative position matrix corresponding to each of the real-time fitness images through the action recognition model; Generate action standard detection information according to the standard feature point relative position matrix corresponding to each standard action image of the target fitness action and the real-time feature point relative position matrix corresponding to each real-time fitness image; Generate fitness action prompt information based on the action standard detection information.