Fitness exercise parameter calculation method and virtual interaction mapping method
By extracting key parameter points from fitness scene images using computer vision technology and combining them with peak detection algorithms to calculate fitness exercise parameters, this approach solves the problems of strong hardware dependence and poor scalability in existing technologies, and achieves low-cost, easy-to-deploy, and universally compatible calculation of fitness exercise parameters.
Patent Information
- Application Number
- CN202511313879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-28
AI Technical Summary
Existing methods for calculating fitness exercise parameters rely on hardware sensors, resulting in high costs, complex installation, poor scalability, and difficulty in adapting to other fitness equipment.
Using computer vision technology, a human key point detection model is used to extract parameter key points from fitness scene images. Combined with a peak detection algorithm, fitness exercise parameters are calculated, and the system is adapted to ordinary cameras for parameter calculation.
It enables low-cost, easy-to-deploy fitness exercise parameter calculation without the need for hardware sensors, and has universal compatibility and agile iteration capabilities, reducing user costs and simplifying the deployment process.
Smart Images

Figure CN120853265A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of fitness exercise parameter detection technology, and particularly relates to fitness exercise parameter calculation methods and virtual interactive mapping methods. Background Technology
[0002] Currently, the interaction between fitness equipment such as exercise bikes, treadmills, rowing machines, and trampolines and virtual scenes largely relies on hardware sensors such as speed sensors, gyroscopes, and pressure sensors to collect user movement data.
[0003] For example, patent CN113076002A discloses an interconnected fitness competition system and method based on multi-body motion recognition. The interconnected fitness competition method includes the following steps:
[0004] (1) Wear the wearable smart unit on the limbs of the human body;
[0005] (2) After the system is started, each wearable smart unit collects the user's posture information in real time, and the acceleration, angular velocity and magnetic field three-axis component data are collected by the accelerometer, gyroscope and magnetic sensor respectively.
[0006] (3) The motion analysis module identifies the quality of the entire motion based on the collected triaxial component data, and generates motion commands based on the relationship between the quality of the motion and the preset motion commands.
[0007] (4) After the action instructions obtained in the above process are sent to the virtual competition module, the animated character can perform the corresponding actions in the virtual scene according to the action instructions. At the same time, various special effects can be added to the virtual scene using virtual engine technology to provide users with various guidance and incentives.
[0008] (5) The fitness data analysis module collects motion quality data during the motion analysis process and competition performance data and fitness-related data during the virtual competition process. After summarizing and analyzing, it obtains user exercise volume and exercise preference data as the basis for fitness data analysis. Combined with the server-side fitness big data and scientific fitness reference data, it provides users with fitness data analysis results and reference suggestions.
[0009] (6) The interconnection module uploads user data to the Internet server. The Internet server aggregates all user data to form fitness big data and feeds it back to the client as reference data. At the same time, the Internet server receives online client interconnection competition data in real time and establishes interconnection channels between multiple clients, so that users can conduct real-time voice and video communication while interconnecting and competing.
[0010] However, these solutions that rely on hardware sensors to collect user motion data have the following drawbacks:
[0011] Strong hardware dependency: Requires additional sensors, increasing costs, installation and maintenance complexity, and thus increasing user operating costs;
[0012] Poor scalability: The solution is designed for specific equipment and is difficult to adapt to other fitness equipment.
[0013] Therefore, there is an urgent need to develop a method for calculating fitness exercise parameters and a virtual interactive mapping method to solve the problems in the existing technology. Summary of the Invention
[0014] The purpose of this invention is to provide a method for calculating fitness exercise parameters and a virtual interactive mapping method, which accurately calculates fitness exercise parameters through computer vision, thereby solving the problems of strong hardware dependence and poor scalability of existing fitness exercise parameter calculation methods mentioned in the background art.
[0015] To solve the above-mentioned technical problems, the specific technical solution of the present invention is as follows:
[0016] A method for calculating fitness exercise parameters includes the following steps:
[0017] Acquire raw images of the user's fitness scene;
[0018] The original images of the fitness scene are input into the human key point detection model to extract human key point data from the original images of the fitness scene.
[0019] Obtain the fitness type and adapt the corresponding parameter calculation algorithm according to the fitness type;
[0020] Based on human body key point data and parameter calculation algorithms, fitness exercise parameters are calculated.
[0021] Furthermore, the human body key point data includes key point coordinate data, and the parameter calculation algorithm includes the following steps:
[0022] Extract the key coordinate data of parameter key points from the key point coordinate data;
[0023] The peak detection algorithm is used to calculate the periodicity of parameters at key points.
[0024] Calculate fitness exercise parameters based on the periodicity of key parameter points.
[0025] Furthermore, the parameter calculation algorithm includes a cadence calculation algorithm, a step frequency calculation algorithm, and a bounce frequency calculation algorithm.
[0026] The key parameters of the cadence calculation algorithm are the key points of the knee joint, the key coordinate data are the vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the knee joint.
[0027] The key parameters of the step frequency calculation algorithm are the key points of the ankle joint, the key coordinate data are horizontal coordinate data and vertical coordinate data, and the parameter periodicity is the periodicity of the horizontal movement of the knee joint.
[0028] The key parameters of the bounce frequency calculation algorithm are the key points of the hip, the key coordinate data are the vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the hip.
[0029] Furthermore, the parameter calculation algorithm also includes a stroke frequency calculation algorithm. The key points of the parameters of the stroke frequency calculation algorithm are the wrist key points and the hip key points, the key coordinate data are horizontal coordinate data, and the parameter periodicity is the periodicity of the horizontal movement of the wrist relative to the hip.
[0030] Furthermore, the human body key point data also includes the confidence level of the key point coordinates, and the parameter key points are key points with high confidence.
[0031] Furthermore, before using the peak detection algorithm to calculate the periodicity of parameters at key points, the following steps are also included:
[0032] Use a moving average filter to remove jitter from key coordinate data.
[0033] Furthermore, the human keypoint detection model includes a preprocessing module, a shallow feature extraction module, a deep feature extraction module, an attention enhancement and multi-scale fusion module, and an output prediction module;
[0034] The process of inputting the original image of the fitness scene into the human key point detection model includes the following:
[0035] The original image of the fitness scene is preprocessed by the preprocessing module to obtain the initial image of the fitness scene;
[0036] The basic features of the initial image of the fitness scene are extracted by the shallow feature extraction module to obtain a shallow feature map;
[0037] The deep feature map is obtained by extracting the deep features from the shallow feature map using the deep feature extraction module.
[0038] The attention enhancement and multi-scale fusion module adjusts the feature channel weights in the deep feature map and fuses multi-scale features to obtain an enhanced pooling feature map.
[0039] The output prediction module maps each spatial location of the feature channel in the enhanced pooling feature map to a key point to obtain human key point data.
[0040] A virtual interaction mapping method includes the following steps:
[0041] Based on the fitness exercise parameter calculation method described above, obtain the user's human body key point data and fitness exercise parameters;
[0042] Control the direction of the virtual scene based on human key point data;
[0043] The virtual character's movements are controlled based on fitness exercise parameters.
[0044] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the fitness exercise parameter calculation method or the virtual interactive mapping method.
[0045] A computer program product includes a computer program that, when executed by a processor, implements the steps of the fitness exercise parameter calculation method or the virtual interactive mapping method.
[0046] The present invention has the following advantages:
[0047] (1) This invention realizes contactless calculation of fitness exercise parameters through computer vision technology, completely eliminating the cost of sensors, controllers and wiring at all equipment end. Only ordinary cameras are required, and there is no need to modify the fitness equipment hardware, which greatly reduces costs and simplifies the deployment process. Users do not need to wear devices or pair them, truly realizing a seamless user experience of "open and use".
[0048] Meanwhile, the solution provided by this invention has universal compatibility. One solution can be adapted to all existing and future fitness equipment without the need for modification of specific equipment. New equipment support can be achieved through algorithm optimization and App updates, and the entire network can be upgraded instantly. It has the characteristics of agile iteration and deployment, providing traditional fitness equipment manufacturers with a shortcut to intelligent upgrades through "AI + software" without the need for hardware R&D capabilities, thus realizing the intelligent empowerment of traditional industries.
[0049] (2) This application ensures the accuracy of fitness exercise parameter calculation by detecting the periodicity of key points of the detection parameters, reduces the requirements of the human body key point detection model, and helps the deployment of the human body key point detection model on mobile devices.
[0050] (3) Use a moving average filter to remove jitter in the key coordinate data to further ensure the accuracy of fitness exercise parameter calculation and reduce the requirements for human key point detection model.
[0051] Other features and advantages of the present invention will be disclosed in detail in the following detailed description and accompanying drawings. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the overall process of this application;
[0053] Figure 2 This is a schematic diagram showing the location of the key human body points in the human body according to this application;
[0054] Figure 3 This is a schematic diagram of the knee joint motion trajectory and peak recognition in this application. Detailed Implementation
[0055] To better understand the purpose, structure, and function of this application, a further detailed description of this application is provided below with reference to the accompanying drawings.
[0056] Example 1
[0057] A method for calculating fitness exercise parameters, such as Figure 1 As shown, it includes the following steps:
[0058] Acquire raw images of the user's fitness scene;
[0059] The original images of the fitness scene are input into the human key point detection model to extract human key point data from the original images of the fitness scene.
[0060] Obtain the fitness type and adapt the corresponding parameter calculation algorithm according to the fitness type;
[0061] Based on human body key point data and parameter calculation algorithms, fitness exercise parameters are calculated.
[0062] Specifically, in this embodiment, acquiring the original image of the user's fitness scene includes the following steps:
[0063] The user places a camera in front of the device, providing a full-body view.
[0064] The camera captures 30 frames per second to obtain raw images of the user's workout scene.
[0065] The original images of the fitness scene can be obtained from images or videos. The types of images or videos include visible light, infrared, or lidar, etc.
[0066] The human keypoint detection model can be an existing human keypoint detection model, used to extract human keypoint data from the original image of a fitness scene.
[0067] like Figure 2 As shown, the human body key points include shoulders, hips, knees, or wrists, etc., and the human body key point data includes the key point coordinate data. Optionally, the human body key point data may also include the confidence level of the key point coordinate data.
[0068] The parameter calculation algorithm is based on the periodic changes of key points on the human body. In this embodiment, the parameter calculation algorithm includes the following steps:
[0069] Extract the key coordinate data of parameter key points from the key point coordinate data;
[0070] The peak detection algorithm is used to calculate the periodicity of parameters at key points.
[0071] Calculate fitness exercise parameters based on the periodicity of key parameter points.
[0072] When the fitness type is spinning, elliptical training, or stepper training, the parameter calculation algorithm is a cadence calculation algorithm. The key parameter of the cadence calculation algorithm is the key point of the knee joint, the key coordinate data is the vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the knee joint.
[0073] Specifically, in this embodiment, the key points of the knee joint are key point 13 and key point 14, and the cadence calculation algorithm specifically includes the following steps:
[0074] Obtain the coordinates of keypoint 13 and keypoint 14 and their corresponding coordinate confidence scores;
[0075] A continuous frame sequence of key points with high coordinate confidence is selected as the main data source, and the vertical coordinate data of the key points in the main data source is extracted. In this embodiment, high coordinate confidence refers to the key point of the knee joint of the leg with less occlusion and clearer motion trajectory in the field of view.
[0076] Use a moving average filter to remove jitter from the Cartesian data;
[0077] The peak detection algorithm is used to calculate the periodicity of vertical knee joint movement;
[0078] The cadence was calculated using an average frequency calculation method based on a time window.
[0079] Specifically, for the coordinate y_i of the i-th frame, its smoothed value y_smooth_i using a moving average filter is calculated using the following formula:
[0080] y_smooth_i=(y_i-2+y_i-1+y_i+y_i+1+y_i+2) / 5.
[0081] When using the PeakDetection algorithm to calculate the periodicity of vertical knee joint movement,
[0082] The height threshold satisfies the following condition: the peak must be higher than the adjacent point by a certain value.
[0083] The distance threshold (distance_threshold) must satisfy the condition that there is at least a 15-frame interval between two peaks, corresponding to a minimum cadence of 30 RPM.
[0084] like Figure 3 As shown, each complete peak + trough cycle corresponds to one complete rotation of the crank of the fitness device.
[0085] When the time window is 5 seconds, the method for calculating the average frequency based on the time window is as follows:
[0086] Count the number of peaks N detected within a 5-second time window; each peak represents half a revolution of the crank, so the total number of revolutions = N / 2.
[0087] The formula for calculating cadence (RPM) is as follows:
[0088] RPM = (Number of laps / Time window in seconds) * 60 = (N / 2) / 5 * 60 = N * 6.
[0089] Example: If 10 peaks are detected within 5 seconds, then RPM = 10 * 6 = 60.
[0090] After the fitness exercise parameters are calculated, the following are also included:
[0091] The calculated values are sent to the application's UI thread to update the display on the screen in real time.
[0092] Optionally, the parameter calculation algorithm further includes the following steps:
[0093] Obtain the motion amplitude of key points based on key coordinate data;
[0094] The motion speed is calculated based on the periodicity of the parameters and the amplitude of motion at the key points of the parameters.
[0095] Specifically, this embodiment also includes:
[0096] The range of motion of the knee joint is obtained from the vertical coordinate data;
[0097] Calculate the movement speed based on cadence and knee joint range of motion.
[0098] When the fitness type is treadmill exercise, the parameter calculation algorithm is a step frequency calculation algorithm. The key points of the step frequency calculation algorithm are the ankle joint key points, and the key coordinate data are horizontal coordinate data and vertical coordinate data. Among them, the vertical coordinate Y is the basis for detecting the stride, but the horizontal coordinate X provides important auxiliary information for verifying and calculating parameters such as stride length and speed. The parameter periodicity is the periodicity of the horizontal movement of the knee joint. In this embodiment, the key points of the ankle joint are key points 15 and 16. The calculation method is the same as above, and will not be repeated in this application.
[0099] When the fitness type is trampoline fitness, the parameter calculation algorithm is a bounce frequency calculation algorithm. The key points of the bounce frequency calculation algorithm are hip key points, the key coordinate data are vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the hip. In this embodiment, the key points of the hip are key point 11 and key point 12, and the displacement of the vertical movement of the hip is calculated by the following formula:
[0100] X = (y 11 +y 12 ) / 2;
[0101] Where X represents the vertical displacement of the hip, and y 11 Let y be the vertical coordinate of key point 11. 12 The vertical coordinates of key point 12.
[0102] These three different calculation algorithms can comprehensively evaluate an athlete's cadence, stride frequency, and bounce frequency during exercise, providing a scientific basis for performance analysis and training guidance. The implementation of these algorithms not only considers the characteristics of different key points but also specifically selects coordinate data and periodic characteristics that best reflect the corresponding motion parameters, thus ensuring the accuracy and practicality of the calculation results.
[0103] When the fitness type is rowing machine exercise, the parameter calculation algorithm is a stroke frequency calculation algorithm, including the following:
[0104] The number of strokes is calculated by counting the number of forward and backward movements of the wrist relative to the hip.
[0105] In this embodiment, the key parameters of the stroke frequency calculation algorithm are wrist key points and hip key points, the key coordinate data are horizontal coordinate data, and the parameter periodicity is the periodicity of the horizontal movement of the wrist relative to the hip. The key points of the wrist are key points 9 and 10, and the key points of the hip are key points 11 and 12.
[0106] In practical applications, this stroke frequency calculation algorithm can accurately calculate the movement cycle by measuring the periodicity of the horizontal movement of the wrist relative to the hip, providing reliable data support for sports performance analysis.
[0107] This application has the following advantages:
[0108] Low cost and easy deployment: Completely eliminates the cost of sensors, controllers and wiring at all equipment end, only ordinary cameras (such as mobile phones) are needed, and there is no need to modify the fitness equipment hardware;
[0109] Universal compatibility: One solution can be adapted to all existing and future fitness equipment without modification;
[0110] Agile iteration and deployment: Algorithm optimization and support for new equipment are all achieved through App updates, completing the entire network upgrade instantly;
[0111] Simple maintenance: Functionality can be optimized by upgrading the software algorithm, without the need for hardware replacement.
[0112] Seamless user experience: Users do not need to wear devices or pair them, truly achieving "open and use".
[0113] Data privacy and security: All processing is completed on the terminal, and sensitive gesture data never leaves the user's device.
[0114] Smart technology empowers traditional industries: providing traditional fitness equipment manufacturers with a shortcut to intelligent upgrades through "AI + software" without requiring hardware R&D capabilities.
[0115] A virtual interaction mapping method includes the following steps:
[0116] Based on the calculation method of fitness exercise parameters, obtain the user's human body key point data and fitness exercise parameters;
[0117] Control the direction of the virtual scene based on human key point data;
[0118] Control the virtual character's movements based on the user's fitness exercise parameters;
[0119] The method of controlling the virtual scene's direction based on human key point data includes the following:
[0120] The virtual scene's rotation and direction are controlled based on the angle of deflection along the line connecting the user's shoulders, i.e., the left and right twisting of the user's body.
[0121] In this embodiment, the deflection angle of the line connecting the two shoulders is the angle between the line connecting the two shoulders in the human body key point data and the horizontal line, wherein the two shoulders are key point 5 and key point 6 respectively.
[0122] The formula for the deflection angle of the line connecting the two shoulders is as follows:
[0123] θ=arctan2(y R -y L ,x R -x L );
[0124] Where θ is the deflection angle of the line connecting the two shoulders, R refers to the key point on the right, L refers to the key point on the left, and x and y are the horizontal and vertical coordinates of the key points, respectively.
[0125] The control of virtual character movements based on the user's fitness exercise parameters includes the following:
[0126] The forward speed of the virtual character is controlled based on the user's cadence, step frequency, or stroke frequency.
[0127] The jump height of the virtual character is controlled based on the user's bounce height.
[0128] Virtual tracks and game levels can dynamically adjust their difficulty and feedback based on fitness exercise parameters.
[0129] Example 2
[0130] In this embodiment, the human keypoint detection model includes a preprocessing module, a shallow feature extraction module, a deep feature extraction module, an attention enhancement and multi-scale fusion module, and an output prediction module.
[0131] The process of inputting the original image of the fitness scene into the human key point detection model includes the following:
[0132] The original image of the fitness scene is preprocessed by the preprocessing module to obtain the initial image of the fitness scene;
[0133] The basic features of the initial image of the fitness scene are extracted by the shallow feature extraction module to obtain a shallow feature map;
[0134] The deep feature map is obtained by extracting the deep features from the shallow feature map using the deep feature extraction module.
[0135] The attention enhancement and multi-scale fusion module adjusts the feature channel weights in the deep feature map and fuses multi-scale features to obtain an enhanced pooling feature map.
[0136] The output prediction module maps each spatial location of the feature channel in the enhanced pooling feature map to a key point to obtain human key point data.
[0137] The preprocessing of the original fitness scene image by the preprocessing module includes the following steps:
[0138] The original image of the input fitness scene is scaled and normalized to obtain the initial image of the fitness scene.
[0139] Specifically, it includes the following steps:
[0140] The original image of the fitness scene was scaled to 192×192 pixels, and the values were converted from the range of [0,255] to the range of [-1,1] to improve training stability and convergence speed.
[0141] In this embodiment, the initial image size of the fitness scene output after normalization is: [1,192,192,3](Batch,Height,Width,Channels).
[0142] The shallow feature extraction module includes a first standard convolutional layer, a rectified linear unit activation function, and a second standard convolutional layer.
[0143] The process of extracting basic features from the initial image of a fitness scene using a shallow feature extraction module includes the following steps:
[0144] The initial image of the fitness scene is convolved by the first standard convolutional layer to obtain the first basic feature map;
[0145] The first basic feature map is processed by the rectified linear unit activation function to obtain the second basic feature map;
[0146] The second basic feature map is convolved by the second standard convolutional layer to obtain a shallow feature map.
[0147] In this embodiment, the first standard convolutional layer uses a set of smaller convolutional kernels to perform convolution operations on the input image to obtain the first basic feature map.
[0148] The Rectified Linear Unit (ReLU) activation function is applied to the first shallow feature map to introduce nonlinearity; in this embodiment, the Rectified Linear Unit activation function sets all negative activation values to zero and retains positive activation values.
[0149] The activation function formula for the rectified linear unit (ReLU) is as follows:
[0150] f(x) = max(0,x).
[0151] The second standard convolutional layer uses a set of smaller convolutional kernels to convolve the input first basic feature map to obtain a shallow feature map.
[0152] In this embodiment, the first standard convolutional layer has 32 output channels, and the second standard convolutional layer has 64 output channels.
[0153] The first standard convolutional layer slides across the image using a set of smaller convolutional kernels to specifically detect basic edges (horizontal, vertical, and diagonal), color patches, simple textures, and corners, ultimately generating a first shallow feature map. Its spatial dimensions may be slightly reduced, and the number of channels increases from 3 to 32, with dimensions approximately [1, 192, 192, 32]. These 32 channels can be understood as 32 different basic feature detectors. Optionally, the spatial dimensions can be preserved through padding.
[0154] The second standard convolutional layer, based on the simple edges and textures extracted by the first standard convolutional layer, combines and abstracts them to begin forming more complex patterns. For example, several edges can be combined into a longer edge, and several intersecting edges may indicate a corner or a more complex texture unit. After passing through the second standard convolutional layer, the number of channels is further increased from 32 to 64, with a size of [1, 192, 192, 64].
[0155] In this embodiment, the depth feature extraction module utilizes the efficiency of depthwise separable convolution to deeply mine intermediate semantic features of the image, such as limb parts and simple shapes, while significantly reducing parameters and computational load. The depth feature extraction module includes at least eight consecutive depthwise separable convolutional blocks, through which depth features in the shallow feature map are extracted.
[0156] The depthwise separable convolutional block includes depthwise convolution, pointwise convolution, batch normalization, and rectified linear unit activation function;
[0157] When the depth-separable convolutional block extracts depth features from the shallow feature map, the following steps are included:
[0158] The shallow feature map is convolved by depthwise convolution to obtain the first depth feature map;
[0159] The second depth feature map is obtained by convolving the first depth feature map with pointwise convolution;
[0160] The third deep feature map is obtained by standardizing the second deep feature map through batch normalization.
[0161] The third depth feature map is processed by the rectified linear unit activation function to obtain the deep feature map.
[0162] Specifically, the depth-separable convolutional block includes:
[0163] A 3×3 kernel depthwise convolution is used to perform spatial convolution on each channel of the input feature map individually;
[0164] Pointwise convolutions with 1×1 kernels are used to combine the outputs of depthwise convolutions;
[0165] Batch normalization is used to accelerate training and improve stability.
[0166] The ReLU activation function is used to introduce nonlinearity.
[0167] In this embodiment, the depthwise convolution is used for spatial filtering, focusing on learning the spatial information within each feature channel. The number of channels in the output feature map is exactly the same as the input, which is 64 channels, and the size may be halved depending on the stride setting.
[0168] The pointwise convolution is used for channel blending, linearly combining features from all channels. It convolves the output of the previous step, creating new feature channels that are weighted combinations of information from all previous channels. This allows the network to learn complex features across channels. For example, it combines information from the "vertical edge" and "skin tone" channels, which might help form the feature of "arm".
[0169] The batch normalization is used to standardize the output of this layer, making its mean close to 0 and its standard deviation close to 1. Batch normalization accelerates the training process, reduces internal covariate bias, allows for higher learning rates, and has a certain regularization effect, making the model more stable.
[0170] This application utilizes the cascading effect of eight depthwise separable convolutional blocks: by stacking these eight blocks, the network can construct a very deep non-linear hierarchical structure. Each subsequent block builds upon the features extracted by the previous block to construct more abstract and semantically meaningful features. From simple edges to textures, and then to parts of limbs (such as forearms and thighs), it is ultimately possible to clearly "see" intermediate components such as joints, head, and torso.
[0171] The attention enhancement and multi-scale fusion module includes an attention enhancement module and a spatial pyramid pooling layer;
[0172] The method of adjusting feature channel weights in deep feature maps and fusing multi-scale features through attention enhancement and multi-scale fusion modules includes the following steps:
[0173] The attention enhancement module obtains the weight of each feature channel in the deep feature map and multiplies the weight back into the deep feature map to obtain the attention-enhanced feature map.
[0174] The attention-enhanced feature map is subjected to multiple max pooling operations at different scales through a spatial pyramid pooling layer. The results of all pooling are then flattened and stitched together to obtain the enhanced pooling feature map.
[0175] The attention enhancement module includes:
[0176] The global average pooling layer (Squeeze) is used to compress the entire feature map (HxW) of each channel into a single value;
[0177] The excitation layer includes a first fully connected layer and a second fully connected layer. The first fully connected layer has ReLU to reduce dimensionality, and the second fully connected layer has Sigmoid to restore the original dimensionality. In this embodiment, the excitation layer can learn a set of weights between 0 and 1, which is the "importance" score of each channel.
[0178] The scaling layer is used to multiply the learned weights back onto the corresponding channels of the feature map, thereby amplifying the responses of the feature channels that are useful in the current pose estimation task and suppressing the responses of unimportant or noisy channels; for example, when recognizing "hands", the weights of channels that are sensitive to hand textures are increased.
[0179] The Spatial Pyramid Pooling (SPP) layer is used to perform multiple max pooling operations at different scales in parallel on the feature map output by the SE module, and flatten and stitch all the pooling results together to capture multi-scale features through pooling kernels of different scales (1×1, 2×2, 4×4).
[0180] In this embodiment, the spatial pyramid pooling layer (SPP) is used to break the fixed-size limitation. Regardless of the size of the human body in the input image, SPP can produce a fixed-length output. It simultaneously captures local details (small pooling kernel) and global context (large pooling kernel) information. This allows the model to effectively detect keypoints regardless of whether the user is near or far from the camera.
[0181] The output prediction module is used to map the high-level features refined through all previous stages to the final prediction—the location of the key point. The output prediction module includes a heatmap head and an offset head.
[0182] The heatmap head includes at least one 1×1 convolution, which acts as a "voter" or "classifier." It maps each spatial location of the deep feature channels to 17 keypoints. The output heatmap is a 3D tensor [H', W', 17], where the value of each point represents the probability that a corresponding keypoint exists at that location.
[0183] The offset head includes at least one 1×1 convolution to predict sub-pixel offsets (dx, dy) at each coarse location in the heatmap. Because the heatmap's resolution is lower than the input image (H' < 192), direct rounding would result in a loss of accuracy. The offset head compensates for this quantization error by regressing precise floating-point offsets, thereby obtaining more accurate coordinates and outputting sub-pixel coordinate offsets to improve positioning accuracy.
[0184] The process of mapping each spatial location of a feature channel in the enhanced pooling feature map to a keypoint via the output prediction module includes the following steps:
[0185] The heatmap header maps each spatial location of the feature channel in the enhanced pooling feature map to a key point, thus obtaining the output heatmap.
[0186] By predicting the sub-pixel-level offset of each coarse location in the heatmap using the offset head, human body key point data is obtained.
[0187] It is understood that this application has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this application. Furthermore, based on the teachings of this application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this application. Therefore, this application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this application.
[0188] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for calculating fitness exercise parameters, characterized in that, Includes the following steps: Acquire raw images of the user's fitness scene; The original images of the fitness scene are input into the human key point detection model to extract human key point data from the original images of the fitness scene. Obtain the fitness type and adapt the corresponding parameter calculation algorithm according to the fitness type; Based on human body key point data and parameter calculation algorithms, fitness exercise parameters are calculated.
2. The method for calculating fitness exercise parameters according to claim 1, characterized in that, The human body key point data includes key point coordinate data, and the parameter calculation algorithm includes the following steps: Extract the key coordinate data of parameter key points from the key point coordinate data; The peak detection algorithm is used to calculate the periodicity of parameters at key points. Calculate fitness exercise parameters based on the periodicity of key parameter points.
3. The method for calculating fitness exercise parameters according to claim 2, characterized in that, The parameter calculation algorithms include a cadence calculation algorithm, a step frequency calculation algorithm, and a bounce frequency calculation algorithm. The key parameters of the cadence calculation algorithm are the key points of the knee joint, the key coordinate data are the vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the knee joint. The key parameters of the step frequency calculation algorithm are the key points of the ankle joint, the key coordinate data are horizontal coordinate data and vertical coordinate data, and the parameter periodicity is the periodicity of the horizontal movement of the knee joint. The key parameters of the bounce frequency calculation algorithm are the key points of the hip, the key coordinate data are the vertical coordinate data, and the parameter periodicity is the periodicity of the vertical movement of the hip.
4. The method for calculating fitness exercise parameters according to claim 3, characterized in that, The parameter calculation algorithm also includes a stroke frequency calculation algorithm. The key points of the parameters of the stroke frequency calculation algorithm are the wrist key points and the hip key points, the key coordinate data are horizontal coordinate data, and the parameter periodicity is the periodicity of the horizontal movement of the wrist relative to the hip.
5. The method for calculating fitness exercise parameters according to claim 3, characterized in that, The human body key point data also includes the confidence level of the key point coordinates, and the parameter key points are key points with high confidence.
6. The method for calculating fitness exercise parameters according to claim 5, characterized in that, Before using the peak detection algorithm to calculate the periodicity of parameters at key points, the following steps are also included: Use a moving average filter to remove jitter from key coordinate data.
7. The method for calculating fitness exercise parameters according to any one of claims 1-6, characterized in that, The human keypoint detection model includes a preprocessing module, a shallow feature extraction module, a deep feature extraction module, an attention enhancement and multi-scale fusion module, and an output prediction module. The process of inputting the original image of the fitness scene into the human key point detection model includes the following: The original image of the fitness scene is preprocessed by the preprocessing module to obtain the initial image of the fitness scene; The basic features of the initial image of the fitness scene are extracted by the shallow feature extraction module to obtain a shallow feature map; The deep feature map is obtained by extracting the deep features from the shallow feature map using the deep feature extraction module. The attention enhancement and multi-scale fusion module adjusts the feature channel weights in the deep feature map and fuses multi-scale features to obtain an enhanced pooling feature map. The output prediction module maps each spatial location of the feature channel in the enhanced pooling feature map to a key point to obtain human key point data.
8. A virtual interactive mapping method, characterized in that, Includes the following steps: According to any one of claims 1-7, the method for calculating fitness exercise parameters obtains the user's human body key point data and fitness exercise parameters. Control the direction of the virtual scene based on human key point data; The virtual character's movements are controlled based on fitness exercise parameters.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the fitness exercise parameter calculation method according to any one of claims 1-7 or the virtual interactive mapping method according to claim 8.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the fitness exercise parameter calculation method according to any one of claims 1-7 or the virtual interactive mapping method according to claim 8.
Citation Information
Patent Citations
Interconnected fitness competition system and method based on multi-part action recognition
CN113076002A