Sports person skeleton point identification method, electronic equipment and storage medium
By jointly training the key point detection model with multi-angle images, the robustness and accuracy issues of skeleton point recognition during motion are solved, and efficient skeleton point recognition is achieved in single-angle images.
Patent Information
- Application Number
- CN202510407596.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-09-19
AI Technical Summary
During exercise, the body position of the athlete changes greatly, which makes the skeleton point recognition prone to self-occlusion. The robustness and accuracy of existing technologies are low.
A key point detection model is used for joint training, and multiple sets of synchronously acquired action images of the same target object from multiple angles and their skeleton point position annotation information are used to improve the robustness and accuracy of the model through joint training operations.
The accuracy and robustness of skeleton point recognition are improved, the computational complexity and hardware requirements are reduced, and accurate recognition is achieved in single-angle images.
Smart Images

Figure CN120673438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and more particularly to a method for identifying skeletal points of an athlete, an electronic device, a storage medium, and a computer program product. Background Art
[0002] In recent years, with the rapid development of science and technology, skeleton point recognition technology has been widely used in fields such as motion analysis, human-computer interaction, and virtual reality. Skeleton point recognition technology aims to identify the key locations of a target object from images or videos and construct a skeleton model. Skeleton points typically refer to the locations of the target object's joints, such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. By identifying these skeleton points, a simplified "skeleton" model can be constructed to represent the target object's posture.
[0003] Related technologies use deep learning models and other techniques to identify the skeleton points of target objects. However, due to the large changes in the body position of the athlete during exercise, self-occlusion often occurs, resulting in large errors in the identification of skeleton points, and low robustness and accuracy. Summary of the Invention
[0004] The present invention has been made in view of the above-mentioned problems.
[0005] According to one aspect of the present invention, a method for identifying skeleton points of an athlete is provided. The method for identifying skeleton points of an athlete comprises:
[0006] Acquiring a first motion image of the athlete;
[0007] Based on a first action image, a key point detection model is used to determine skeletal point information of an athlete in the first action image, wherein the key point detection model is obtained by performing a joint training operation using multiple groups of second action images and position annotation information of skeletal points in each second action image, and each group of second action images includes second action images of the same target object acquired synchronously from multiple angles.
[0008] Exemplarily, before determining the skeleton point information of the athlete in the first action image using a key point detection model based on the first action image, the method for identifying the skeleton points of the athlete also includes: determining the position prediction information of the athlete in the first action image using a target detection model based on the first action image; processing the first action image according to the position prediction information of the athlete in the first action image to obtain a first target image corresponding to the first action image, wherein the first target image includes the area where the athlete is located; determining the skeleton points of the athlete in the first action image using a key point detection model based on the first action image, including: inputting the first target image into the key point detection model to obtain the skeleton point information of the athlete.
[0009] Exemplarily, the target detection model is obtained by performing a joint training operation using multiple groups of third action images and position annotation information of the athletes in each third action image, wherein each group of third action images includes third action images of the same target object acquired synchronously from multiple angles.
[0010] Exemplarily, the joint training operation includes: for a second action image at another angle that is acquired synchronously with a second action image at a first angle, converting the position annotation information of the skeleton points in the second action image into position information at the first angle as auxiliary supervision data of the second action image at the first angle, wherein the first angle is one of multiple angles; inputting the second action image at the first angle into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the second action image at the first angle; training the key point detection model according to a first loss function value, wherein the first loss function value is calculated based on at least the position prediction information, the first main supervision data and the auxiliary supervision data of the second action image at the first angle, and the first main supervision data includes the position annotation information of the skeleton points of the athlete in the second action image at the first angle.
[0011] Exemplarily, the joint training operation also includes: for each second angle among multiple angles that is different from the first angle, converting the position annotation information of the skeleton points in the second action images of angles other than the second angle, which are obtained synchronously with the second action image of the second angle, into position information at the second angle as auxiliary supervision data of the second action image of the second angle; inputting the second action image of the second angle into the key point detection model to obtain the position prediction information of the skeleton points of the athlete in the second action image of the second angle; wherein the first loss function value is also calculated based on the position prediction information of the skeleton points of the athlete in the second action image of the second angle, the second main supervision data and the auxiliary supervision data of the second action image of the second angle, and the second main supervision data includes the position annotation information of the skeleton points of the athlete in the second action image of the second angle.
[0012] Exemplarily, the joint training operation includes: inputting at least the first image in the second action image into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the first image; determining the polar line corresponding to each skeleton point in the first image based on the basic matrix corresponding to the first image and the second image in the second action image, wherein the second image is an image acquired synchronously with the first image at a different angle, and the basic matrix represents the geometric relationship between the first image and the second image; training the key point detection model according to the second loss function value, wherein the second loss function value is calculated based on the distance from the position prediction information of each skeleton point in the first image to its corresponding polar line.
[0013] Exemplarily, the joint training operation includes: inputting at least one second action image into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the second action image; training the key point detection model according to a third loss function value, wherein the third loss function value is based on the position prediction information of the skeleton points in the second action image, the position annotation information of the skeleton points in the second action image and other second action images, and the skeleton point weight calculation, wherein, among all the skeleton points in the second action image, there are at least two skeleton points with different skeleton point weights.
[0014] According to another aspect of the present invention, an electronic device is provided, comprising: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, they are used to execute the above-mentioned method for identifying skeleton points of an athlete.
[0015] According to another aspect of the present invention, a storage medium is provided, on which program instructions are stored. When the program instructions are run, they are used to execute the above-mentioned method for identifying skeleton points of an athlete.
[0016] According to another aspect of the present invention, a computer program product is provided, comprising computer program instructions, which are used to execute the above-mentioned method for identifying skeleton points of an athlete when the computer program instructions are run.
[0017] The above technical solution determines the skeleton points of the athlete based on the first action image using a key point detection model. The key point detection model is obtained by performing a joint training operation using the second action images acquired synchronously from multiple angles and the position annotation information of the skeleton points in each of the second action images. The key point detection model is jointly trained using the second action images from multiple angles, so that the model can learn more skeleton point features, compensate for the information deficiencies caused by the athlete's self-occlusion or rapid movement, and improve the accuracy and robustness of identifying skeleton points. Moreover, during recognition, only a single-angle first action image is required for accurate skeleton point recognition, without the need for images from multiple angles, thereby reducing the computational complexity and the hardware requirements in practical applications.
[0018] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other objects, features, and advantages of the present invention will become more apparent through a more detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings are provided to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and are not intended to limit the present invention.
[0020] Figure 1 A schematic flow chart of a method for identifying skeleton points of an athlete according to an embodiment of the present application is shown;
[0021] Figure 2 A schematic diagram showing detected skeleton points of an athlete according to one embodiment of the present application is shown;
[0022] Figure 3 A schematic diagram showing lines connecting skeletal points of an athlete in a first action image according to an embodiment of the present application is shown;
[0023] Figure 4 FIG2 shows a schematic flow chart of a joint training operation according to an embodiment of the present application;
[0024] Figure 5 shows a schematic flow chart of a joint training operation according to another embodiment of the present application;
[0025] Figure 6 FIG2 shows a schematic flow chart of a joint training operation according to another embodiment of the present application;
[0026] Figure 7 A schematic block diagram of a device for identifying skeleton points of an athlete according to an embodiment of the present application is shown;
[0027] Figure 8 A schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more obvious, the exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention.
[0029] In recent years, significant progress has been made in AI-based research on computer vision, deep learning, machine learning, image processing, and image recognition. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. AI is a comprehensive discipline encompassing numerous technologies, including chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. Computer vision, a key branch of AI, specifically enables machines to understand the world. Computer vision technologies typically include face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, and robotic navigation and positioning. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, autonomous driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, smart temperature measurement and other fields.
[0030] In the training process of existing artificial intelligence models used for skeleton point recognition methods, the AI model is typically trained separately using each training image in the training image set. However, in many sports such as tennis and gymnastics, skeleton point recognition methods using these AI models face challenges in recognition accuracy due to the rapid movement of athletes and frequent self-occlusion (such as the torso occluding the elbow).
[0031] In this application, the acquisition and use of all image data have been explained to the user in advance regarding purpose, scope of use, etc., and have been agreed to by the user to fully protect the user's privacy.
[0032] In order to at least solve the above technical problems, the present application proposes a method for identifying skeleton points of an athlete. Figure 1 FIG. 1 is a schematic flow chart of a method for identifying the skeleton points of an athlete according to an embodiment of the present application. Figure 1 As shown, the method for identifying skeleton points of an athlete includes the following steps S1100 and S1200.
[0033] In step S1100 , a first motion image of an athlete is acquired.
[0034] In step S1200, based on the first action image, a key point detection model is used to determine skeletal point information of the athlete in the first action image. The key point detection model is obtained by performing a joint training operation using multiple sets of second action images and the positional annotation information of the skeletal points in each second action image. Each set of second action images includes synchronously acquired second action images of the same target object from multiple angles.
[0035] Athletes may include athletes participating in competitions or training, patients undergoing rehabilitation training, and the like. The first motion image is an image captured during the athlete's movement. Optionally, in step S1100, an image or video of the athlete may be captured using a camera device to obtain the first motion image. For example, if the video frame rate is 30 frames per second, 30 frames of image may be obtained in 1 second, and all or part of the images in the video may be used as the first motion image. Alternatively, the first motion image of the athlete may be obtained from a storage device stored therein.
[0036] In step S1200, the key point detection model is used to determine the skeleton point information of the athlete in the first action image.
[0037] Exemplarily, the skeletal point information may represent the positions of the skeletal points detected by the athlete, for example, the skeletal point information may include the coordinates, identifications, etc. of the skeletal points. First, a preprocessing operation may be performed on the first action image, and then the preprocessed first action image may be input into a key point detection model to determine the skeletal point information of the athlete in the first action image. The preprocessing operation may include all operations for improving the visual effect of the first action image, increasing its clarity, or highlighting certain features in the image. Exemplarily and not restrictively, the preprocessing operation may include digitization, geometric transformation, normalization, filtering, and other operations on the first action image.
[0038] Users can set the skeletal points of an athlete to be detected by the key point detection model based on their detection requirements. For example, you can set the skeletal points of an athlete to be detected based on the detection accuracy requirements. If the accuracy requirement is high, more skeletal points can be detected; if the accuracy requirement is low, fewer skeletal points can be detected.
[0039] Figure 2 FIG. 1 shows a schematic diagram of detected skeleton points of an athlete according to an embodiment of the present application. Figure 2 As shown, the detected bone points include the top of the head 1 (the center of the top of the head, which is also the highest point of the human body), the neck 2 (the center point of the neck), the left wrist point 3 (the center point of the left wrist joint), the left elbow point 4 (the center point of the left elbow joint), the left shoulder point 5 (the center point where the left upper arm is connected to the human body), the right wrist point 6 (the center point of the right wrist joint), the right elbow point 7 (the center point of the right elbow joint), the right shoulder point 8 (the center point where the right upper arm is connected to the human body), the left hip outer contour 9 (the outer contour point of the left hip), the left knee point 10 (the left knee joint). Joint center point), left ankle point 11 (left ankle joint center point), right hip outer contour 12 (right hip outer contour point), right knee 13 (right leg knee joint center point), right ankle point 14 (right ankle joint center point), left chest 15 (left chest midpoint), right chest 16 (right chest midpoint), left waist 17 (left waist outer point), waist center 18 (waist center point), right waist 19 (right waist outer point), left thigh starting point 20 (left thigh outer starting point), crotch 21 (center point of both legs starting point), right thigh starting point 22 (right thigh outer starting point).
[0040] The key point detection model used in step S1200 is obtained by performing a joint training operation using multiple groups of second motion images and the position annotation information of the skeleton points in each second motion image. Each group of second motion images includes second motion images of multiple angles of the same target object acquired synchronously. It can be understood that the joint training here refers to training the key point detection model by integrating the second motion images from multiple angles and the position annotation information in each of them. Therefore, even if the key point detection model is used to process the first motion images collected in self-occlusion and dynamic scenes, it will not perform poorly due to the limitations of the single shooting angle of the first motion image. The key point detection model can include any existing or future developed model or algorithm that can be used to detect key points in an image. For example, the key point detection model can be implemented using HRNet, AlphaPose, etc.
[0041] The target object can be any athlete. The second action image in different groups can be shot for different target objects. The second action image in the same group can be shot synchronously at multiple angles for the same target object. Multiple cameras can be utilized to synchronously shoot videos or images of any athlete at different angles. For example, three cameras can be adopted, with the athlete being shot as the center, and the athlete is shot synchronously at 0 degree, 120 degree and 240 degree respectively, using each frame image in the video as the second action image. The frame rates of the three cameras are the same. In other words, the second action images of the athlete of three different viewing angles can be obtained at the same time point. The skeleton points of the athlete in the second action image can be marked to obtain the position marking information of the skeleton points corresponding to each second action image. It is understandable that the skeleton points marked can be identical with the skeleton points detected above. The skeleton points marked can be more than the skeleton points detected above. In other words, the skeleton points marked at least include the skeleton points detected above.
[0042] In some embodiments, an image from one of the multiple synchronously acquired second motion images can be input into a key point detection model and its prediction result can be obtained. Then, based on the position annotation information of the skeleton points of the second motion image from that angle and the second motion images from the remaining angles and the prediction result, the parameters of the key point detection model can be adjusted to achieve joint training of the key point detection model. In other embodiments, the multiple synchronously acquired second motion images can all be input into a key point detection model to obtain the prediction result. Then, based on the prediction result and the position annotation information of the skeleton points of the second motion image from each angle, the key point detection model can be jointly trained. In these embodiments, the second motion images from multiple angles can supervise each other.
[0043] After jointly training the key point detection model with training images from multiple angles, the key point detection model, which meets the required performance requirements, can be deployed to determine the skeletal points of the athlete in the first action image. The first action image is then input into the key point detection model to determine the skeletal points of the athlete in the first action image. The athlete's skeletal points can be used to reconstruct the athlete in three dimensions. Accurate data on the athlete's skeletal points can improve the quality of the 3D reconstruction. Consequently, the athlete's skeletal points can be used for motion analysis, virtual and augmented reality, the gaming industry, and more.
[0044] The above technical solution determines the skeleton points of the athlete based on the first action image using a key point detection model. The key point detection model is obtained by performing a joint training operation using the second action images acquired synchronously from multiple angles and the position annotation information of the skeleton points in each of the second action images. The key point detection model is jointly trained using the second action images from multiple angles, so that the model can learn more skeleton point features, compensate for the information deficiencies caused by the athlete's self-occlusion or rapid movement, and improve the accuracy and robustness of identifying skeleton points. Moreover, during recognition, only a single-angle first action image is required for accurate skeleton point recognition, without the need for images from multiple angles, thereby reducing the computational complexity and the hardware requirements in practical applications.
[0045] Exemplarily, the method for identifying skeleton points of an athlete further includes step S1300. In step S1300, the action of the athlete in the first action image is classified and / or evaluated based on the skeleton points of the athlete in the first action image.
[0046] After acquiring the skeleton points of the athlete in the first motion image, corresponding skeleton points are connected by lines, and a general posture of the athlete can be basically determined based on the angular relationship between the lines connecting any two skeleton points. Figure 3 FIG. 1 is a schematic diagram showing the lines connecting the skeleton points of the athlete in the first action image according to an embodiment of the present application. Figure 3 As shown, there are 18 skeleton points in total in the first motion image. The top of the head 0 is connected to the neck 1, the neck 1 is respectively connected to the left shoulder 2 and the right shoulder 3, the left shoulder 2 is connected to the left elbow 4, the left elbow 4 is connected to the left wrist 6, the right shoulder 3 is connected to the right elbow 5, the right elbow 5 is connected to the right wrist 7, the left shoulder 2 is connected to the left hip 8, the left hip 8 is connected to the left knee 10, the left knee 10 is connected to the left ankle 12, the left ankle 12 is connected to the left heel 14, the left heel 14 is connected to the left toe 16, the right shoulder 3 is connected to the right hip 9, the right hip 9 is connected to the right knee 11, the right knee 11 is connected to the right ankle 13, the right ankle 13 is connected to the right heel 15, and the right heel 15 is connected to the right toe 17. The connection lines of the skeleton points in the first motion image can roughly determine the posture of the athlete, and the posture of the athlete can be further judged according to the specific sports scene performed by the athlete. For example, Figure 3 The figure shows the skeleton point connection lines in the first motion image captured when the athlete is doing tennis training, based on which it can be determined that the athlete's action at this time is a forehand shot.
[0047] In some embodiments, the skeleton point information of the athlete in the first motion image can be input into a classification model, and the classification model can automatically classify the action of the athlete in the first motion image based on the skeleton point information of the athlete in the first motion image. The classification model may include models such as PoseC3D and PoTion. The classification model may be pre-trained and can identify specific actions in different sports scenes. For example, the athlete can play tennis, and the classification model is pre-trained to classify the actions of tennis. For example, the actions of tennis may include: serving, forehand strokes, backhand strokes, high-pressure balls, slices, and net volleys. The classification model can determine the action of the athlete as one of the above six actions based on the skeleton points of the athlete in the first motion image.
[0048] Exemplarily, the athlete's action can be evaluated based on the first action image and the corresponding standard action template. It is understandable that a sports action can last for a period of time, and the action evaluation can be performed based on multiple continuous first action images and the corresponding standard action template. The action template can be multiple continuous images of a sports action. By comparing the difference between the first action image and the standard action image, the athlete's action can be evaluated. Preferably, the dynamic time warping (DTW) algorithm can be used to evaluate the classified action. The DTW algorithm is a method for measuring the similarity between two time series, which is applicable to time series data. It has time elasticity and can allow the number of multiple continuous first action images to be inconsistent with the number of standard action images, that is, the two can be misaligned in time. In other words, the difference in action rhythm can be handled better.
[0049] Preferably, first, the athlete's movements are classified based on the athlete's skeletal points in the first movement image, i.e., the athlete's movements in the first movement image are determined; then, the athlete's movements are evaluated. Thus, when evaluating the athlete's movements, the evaluation can be performed on the first movement image in which the athlete has the same movement, further improving the accuracy of the movement evaluation.
[0050] After acquiring the first action image and determining the athlete's skeletal points, action classification and evaluation can be performed in real time. In other words, each time an athlete's skeletal points are obtained in the first action image, the athlete's action classification and evaluation can be performed. The aforementioned key point detection model, classification model, and evaluation algorithm can be combined into a hybrid model to rapidly determine the athlete's skeletal points in the first action image and subsequently perform action classification and evaluation. This improves efficiency and real-time performance.
[0051] The above technical solution can classify and / or evaluate the athlete's movements based on the athlete's skeletal points in the first action image. This can help better evaluate the athlete's training effect and training intensity, and improve the effectiveness and accuracy of sports training and medical rehabilitation.
[0052] For example, Figure 4 FIG. 1 shows a schematic flow chart of a joint training operation according to an embodiment of the present application. Figure 4 As shown, the joint training operation includes: step S4100, step S4200 and step S4300. The key point detection model is trained through the joint training operation to obtain the key point detection model for determining the skeleton points in the first action image.
[0053] In step S4100, for a second action image at another angle acquired synchronously with the second action image at the first angle, the position annotation information of the skeletal points in the second action image is converted into position information at the first angle to serve as auxiliary supervision data for the second action image at the first angle, where the first angle is one of the multiple angles.
[0054] For example, there may be second motion images captured synchronously from three different angles: 0 degrees, 120 degrees, and 240 degrees. The first angle may be any one of these three angles, for example, 0 degrees. The position annotation information of the skeletal points in the 120-degree and 240-degree second motion images may be converted to position information at 0 degrees using a projection transformation matrix, serving as auxiliary supervision data for the 0-degree second motion image. The projection transformation matrix may be determined based on the spatial relationship between the camera device at the first angle and the other angles. Still using the three aforementioned angles as an example, the position annotation information for the second motion image corresponding to 120 degrees can be converted to position information for the second motion image corresponding to 0 degrees based on the relative positional relationship between the camera device at 120 degrees and the camera device at 0 degrees, serving as auxiliary supervision data for the second motion image at the first angle (0 degrees). Similarly, the position annotation information for the 240-degree second motion image may also be converted to position information for the second motion image at the first angle (0 degrees), serving as auxiliary supervision data for the second motion image at the first angle (0 degrees).
[0055] For example, multiple second motion images from different angles can be synchronously acquired at a certain frequency, and these synchronously acquired second motion images from different angles can be considered a group of images. In other words, multiple groups of images can be acquired sequentially, and the annotation information of the second motion images from other angles in each group of images, after being converted into position information at the first angle, can be used as auxiliary supervision data for the second motion images from the first angle in the group of images.
[0056] In step S4200, the second action image at the first angle is input into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the second action image at the first angle.
[0057] In a set of images that includes multiple second motion images acquired simultaneously from different angles, the second motion images from the first angle in each set of images can be input into a key point detection model, which then outputs predicted information about the positions of the skeletal points of the athlete in each second motion image from the first angle. In other words, the second motion images from other angles can be omitted from the key point detection model, and only the second motion images from the first angle can be input into the key point detection model for skeletal point recognition.
[0058] In step S4300, the key point detection model is trained according to the first loss function value, wherein the first loss function value is calculated based on at least position prediction information, first main supervision data and auxiliary supervision data of the second action image at the first angle, and the first main supervision data includes position annotation information of the skeleton points of the athlete in the second action image at the first angle.
[0059] The error of the key point detection model can be determined based on the first loss function value, and the parameters of the key point detection model can be adjusted based on the first loss function value until preset conditions are met, thereby completing the training. The preset conditions may include that the model error meets preset requirements, the number of training times reaches a preset number, etc. The trained key point detection model can then be deployed as the key point detection model used in step S1200 to determine the skeletal points of the athlete in the first action image.
[0060] The first primary supervisory data includes the positional annotation information of the athlete's skeletal points in the second action image from the first angle. In other words, the positional annotation information in the second action image from the first angle serves as the first primary supervisory data. The positional information converted from the positional annotation information of the skeletal points in the second action images from other angles serves as the auxiliary supervisory data.
[0061] The first loss function can be in the form of an absolute value error loss function, a mean square error loss function, or the like. The deviations between the position prediction information of each skeletal point in the second motion image at each first angle and the first primary supervisory data and the auxiliary supervisory data can be calculated, respectively, to thereby calculate the value of the first loss function. For example, if the first loss function is an absolute value error loss function, the absolute value errors between the position prediction information of each skeletal point in the second motion image at each first angle and the first primary supervisory data and the auxiliary supervisory data can be calculated, respectively, to thereby calculate the value of the first loss function.
[0062] Optionally, weights can be set in the first loss function. For example, the deviation between the position prediction information and the first primary supervisory data can be given a larger weight, while the deviation between the position prediction information and the auxiliary supervisory data can be given a smaller weight. In other words, the former is more important than the latter. The weights in the first loss function can be adjusted according to the needs of the joint training operation.
[0063] The above technical solution converts the position annotation information of the skeletal points in the second action image at angles other than the first angle into position information at the first angle, which serves as auxiliary supervision data for the second action image at the first angle. The first loss function value is calculated based on the position prediction information of the second action image at the first angle, the first primary supervision data, and the auxiliary supervision data to train the key point detection model. By adding auxiliary supervision data from other angles to supervise the training of the key point detection model, the key point detection model's ability to detect the skeletal points of athletes at different angles is improved. Multi-angle supervision can compensate for the information loss caused by the athlete's self-occlusion, thereby improving the accuracy and robustness of the key point detection model.
[0064] Exemplarily, the joint training operation further includes: executing step S4400 and step S4500 for each second angle different from the first angle among the multiple angles.
[0065] In step S4400, the position annotation information of the skeleton points in the second action image at an angle other than the second angle, which is obtained synchronously with the second action image at the second angle, is converted into position information at the second angle to serve as auxiliary supervision data for the second action image at the second angle.
[0066] Still taking the above-mentioned multiple angles including 0 degrees, 120 degrees, and 240 degrees as an example, the second angle can include all or part of the other angles other than the first angle. If the first angle is 0 degrees, then 120 degrees can be the second angle, and 240 degrees can also be the second angle. For example, taking the second angle of 120 degrees as an example, the position annotation information of the skeleton points in the second action image corresponding to 0 degrees and 240 degrees can be converted into position information at 120 degrees to serve as auxiliary supervision data for the second action image of 120 degrees. The conversion method can be the same as the conversion method in the above-mentioned step S4100. For the sake of brevity, it will not be repeated here. In this way, the auxiliary supervision data of the second action image at each second angle can be determined.
[0067] In step S4500, the second action image at the second angle is input into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the second action image at the second angle.
[0068] Similarly, each second action image corresponding to the second angle can be input into the key point detection model, so that the key point detection model outputs the position prediction information of the skeleton points of the athlete in each second action image at the second angle.
[0069] It is understood that step S4400 and step S4500 can be performed before step S4300. When the key point detection model is trained according to the first loss function value in step S4300, the first loss function value is also calculated based on the position prediction information of the skeletal points of the athlete in the second action image at the second angle, the second primary supervisory data, and the auxiliary supervisory data of the second action image at the second angle, where the second primary supervisory data includes the position annotation information of the skeletal points of the athlete in the second action image at the second angle.
[0070] The second primary supervisory data is positional annotation information of the skeletal points of the athlete in the second motion image at the second angle. When calculating the value of the first loss function, the deviation between the position prediction information of each skeletal point in each second motion image at the first angle and the first primary supervisory data and the corresponding auxiliary supervisory data can be calculated, and the deviation between the position prediction information of each skeletal point in each second motion image at the second angle and the second primary supervisory data and the corresponding auxiliary supervisory data can be calculated, thereby calculating the value of the first loss function.
[0071] In short, in this technical solution, the second motion images at different angles are input into the key point detection model. For each second motion image input into the key point detection model, its corresponding position prediction information can be obtained, and the first loss function value is calculated based on the deviation between the position prediction information and the corresponding primary supervision data and auxiliary supervision data. The second motion images input into the key point detection model can be at least two images from the second motion images at multiple angles acquired synchronously. Preferably, all the second motion images at multiple angles can be input into the key point detection model, and the corresponding first loss function values can be calculated to jointly train the key point detection model.
[0072] The above technical solution inputs the second action image from multiple angles into the key point detection model to obtain corresponding position prediction information. Based on the position prediction information, the corresponding primary and auxiliary supervisory data, the first loss function value is calculated to train the key point detection model. This allows the key point detection model to learn the characteristics of the same action performed by the athlete at different angles, further improving the accuracy and robustness of the key point detection model.
[0073] For example, Figure 5 FIG. 1 shows a schematic flow chart of a joint training operation according to another embodiment of the present application. Figure 5As shown, the joint training operation includes step S5100, step S5200 and step S5300.
[0074] In step S5100, at least the first image in the second action image is input into a key point detection model to obtain position prediction information of the skeleton points of the athlete in the first image.
[0075] In some embodiments, the second motion images at one or more specific angles among the second motion images at multiple angles acquired synchronously may be used as the first image. In other embodiments, for the sake of simplicity, all the second motion images acquired at multiple angles simultaneously may be regarded as a group of images, and multiple groups of images may be acquired continuously. The first image may include any one or more images in a group of images. In other words, the first image may not be an image at a specific angle, but may be a random image in a group of images. The first image may be determined according to a specific rule among the multiple groups of images acquired continuously. For example, multiple groups of second motion images are included, and each group of second motion images includes second motion images at three angles. In the first group, the first image may be the second motion image at angle 1; in the second group, the first image may be the second motion image at angle 2; in the third group, the first image may be the second motion image at angle 3; in the fourth group, the first image may again be the second motion image at angle 1...
[0076] In step S5200, the epipolar line corresponding to each skeletal point in the first image is determined based on the fundamental matrix corresponding to the second image in the first and second action images. The second image is an image acquired synchronously with the first image at a different angle, and the fundamental matrix represents the geometric relationship between the first and second images.
[0077] The second image is an image acquired synchronously with the first image and at a different angle than the first image. In other words, the second image and the first image are in the same set of second motion images. Still taking a set of second motion images comprising three angles as an example, the first image can be a second motion image at angle 1, the second image can be a second motion image at angle 2, and the second image can also be a second motion image at angle 3. The fundamental matrix represents the geometric relationship between the first and second images, and can represent the positional correspondence between pixels at the same point on the athlete captured in the first and second images. The fundamental matrix can be determined based on the intrinsic parameter matrix of the camera device and the positions at which the camera device captured the first and second images, respectively. In other words, two second motion images at different angles correspond to one fundamental matrix. The epipolar line corresponding to each skeletal point in the first image can be determined based on the fundamental matrix corresponding to both the first and second images and the annotated position information of each skeletal point.
[0078] In step S5300, the key point detection model is trained according to a second loss function value, wherein the second loss function value is calculated based on the distance between the position prediction information of each skeleton point in the first image and its corresponding epipolar line.
[0079] In theory, the predicted position information for each skeletal point in the first image should lie on the corresponding epipolar line. If it deviates from the epipolar line, an error occurs. A second loss function value can be calculated based on the distance between the predicted position information for each skeletal point in the first image and its corresponding epipolar line. The second loss function can take the form of an absolute value error loss function, a mean square error loss function, or the like.
[0080] Exemplarily, the second loss function value can be calculated according to the following formula:
[0081]
[0082] Among them, L epipolar Represents the second loss function value; i and j represent the first image and the second image at different angles, respectively. ij 、b ij 、c ij Respectively represent the corresponding polar parameters; (x B ,y B ) represents the position coordinates of the skeleton points in the first image. Thus, the distance from each skeleton point to the corresponding epipolar line can be calculated, and the sum of the distances can be used as the second loss function.
[0083] The first image and the second image can be shot at different angles. Different calculation combinations can be formed in the second action images of multiple angles acquired simultaneously. Taking three angles as an example, the shooting angle of the first image can be angle 1, which can form a calculation combination with the second action image of angle 2, and the distance from the position prediction information of the skeleton point to the corresponding polar line is calculated in each of them. Similarly, the second action image of angle 1 can also be used as the first image and form a calculation combination with the second action image of angle 3. Of course, the second action image of angle 2 can also be used as the first image and form a calculation combination with the second action image of angle 3. For the second action images of multiple angles acquired simultaneously, the user can determine the calculation combination and then input the second action images that need to be calculated into the key point detection model to obtain their corresponding skeleton point position prediction information. The second action images that do not participate in the calculation can be omitted from the key point detection model.
[0084] The above technical solution inputs the first image into the keypoint detection model, obtains the position prediction information of the corresponding skeletal points, and then determines the epipolar line corresponding to each skeletal point in the first image based on the fundamental matrices corresponding to the first and second images. The keypoint detection model is then trained by calculating the value of a second loss function based on the distance between the position prediction information and the corresponding epipolar line. This allows the consistency of the keypoint detection model's detection from multiple angles to be measured, thereby improving the consistency of its detection of people in different postures and enhancing the robustness of the keypoint detection model.
[0085] Although the above-described various embodiments describe training the key point detection model based on the first loss function value and the second loss function value, it is understood that the key point detection model can be trained using both the first loss function value and the second loss function value. Thus, the trained key point detection model can detect skeletal points in the first action image from different angles more accurately and more consistently. In other words, the skeletal points of the athlete in the first action image determined using the key point detection model will be more ideal.
[0086] For example, Figure 6 FIG. 4 shows a schematic flow chart of a joint training operation according to another embodiment of the present application. Figure 6 As shown, the joint training operation includes step S6100 and step S6200.
[0087] In step S6100, at least one second action image is input into a key point detection model to obtain position prediction information of skeleton points in the second action image.
[0088] This step S6100 is similar to the aforementioned step S4200 and step S5100 and will not be described again for the sake of brevity.
[0089] In step S6200, a key point detection model is trained based on a third loss function value, wherein the third loss function value is calculated based on the position prediction information of the skeletal points in the second action image, the position annotation information of the skeletal points in the second action image and other second action images, and the weights of the skeletal points. Each skeletal point weight is applied to skeletal points in different parts of the athlete. Among all the skeletal points of the athlete in the second action image, there are at least two skeletal points with different weights.
[0090] For each second action image, a third loss function value can be calculated based on the deviation between the predicted position information of each skeletal point in the second action image and the position annotation information of the corresponding skeletal point in the second action image and other second action images. Furthermore, the third loss function value can be calculated based on the deviation of each skeletal point and the corresponding skeletal point weight. In other words, the impact of the deviation of each skeletal point on the third loss function value can be different. The larger the skeletal point weight, the greater the impact of the deviation of the skeletal point on the third loss function value. For example, the skeletal point weight of each skeletal point can be determined based on the different body parts of the athlete where the skeletal point is predicted to be located or its different importance to the sport. For example, skeletal points on the head can be assigned a larger weight because the parts corresponding to these skeletal points are smaller and the deviation has a greater impact. Skeletal points on the hips can be assigned a smaller weight because the parts corresponding to these skeletal points are larger and the deviation has a smaller impact. In other embodiments, the skeletal point weights can be obtained by statistically analyzing existing skeletal point recognition datasets. For example, statistical calculations can be performed in the COCO dataset. Specifically, the skeletal point weights of different skeletal points can be determined by statistically analyzing the object keypoint similarity (OKS) of different skeletal points in different images.
[0091] Optionally, when calculating the first loss function value and / or the second loss function value, bone point weights may be set for the skeleton points to achieve different effects of the deviations of different skeleton points on the loss function value. In other words, the third loss function value in this embodiment may include the first loss function value and / or the second loss function value.
[0092] The above technical solution calculates a third loss function value based on the position prediction information of the skeletal points in the second action image, the position annotation information of the skeletal points in the second action image and other second action images, and the weights of the skeletal points to train the key point detection model. The deviations of different skeletal points have different effects on the loss function value. Important skeletal points can be assigned larger skeletal point weights, while unimportant skeletal points can be assigned smaller skeletal point weights. This can improve the key point detection model's ability to detect important skeletal points, thereby improving the overall recognition accuracy of the athlete's skeletal points, making the results more consistent with expectations.
[0093] Exemplarily, the method for identifying skeleton points of an athlete further includes step S1400 and step S1500. Both step S1400 and step S1500 are performed before step S1200 of determining the skeleton points of the athlete in the first action image.
[0094] In step S1400 , based on the first action image, the target detection model is used to determine the position prediction information of the athlete in the first action image.
[0095] The first action image may include both the area where the athlete is located and the surrounding background area. In other words, the area where the athlete is located is part of the first action image. The object detection model can be used to determine the athlete's position prediction information in the first action image to determine the athlete's specific position in the first action image.
[0096] The target detection model can be implemented using any existing or future target detection model. For example, the target detection model can be implemented using the YOLOv5 model. Exemplarily, the position prediction information can be in the form of a bounding box of the athlete. After determining the area where the athlete is located in the first action image, the target detection model can set a bounding box around the area where the athlete is located to serve as the position prediction information of the athlete. The bounding box can be, for example, a minimum circumscribed rectangular box. Alternatively, the position prediction information can also be the outline of the athlete, that is, the position prediction information includes edge information of the athlete.
[0097] In step S1500, the first action image is processed based on the predicted position information of the athlete in the first action image to obtain a first target image corresponding to the first action image. The first target image includes the area where the athlete is located. The predicted position information indicates the athlete's position in the first action image.
[0098] In some embodiments, the first action image can be cropped based on the position prediction information of the athlete in the first action image. For example, cropping can be performed along the corresponding bounding box in the position prediction information, and the inside of the bounding box is the retained image, that is, the first target image. It can be understood that the inside of the bounding box is the area where the athlete is located, and the first target image includes the inside of the bounding box, that is, the first target image includes the area where the athlete is located. The first target image includes the area where the athlete is located and may also include some surrounding background areas. Optionally, the bounding box corresponding to the position prediction information can be extended outward or contracted inward by a certain distance before cropping.
[0099] In other embodiments, the first motion image may be cropped based on the predicted position of the athlete in the first motion image. For example, after determining the athlete's position in the first motion image based on the predicted position information, the first motion image may be cropped along the athlete's outline, with the cropped image serving as the first target image. The first target image may include only the area where the athlete is located, excluding the background area surrounding the athlete in the first motion image.
[0100] Step S1200, based on a first action image, uses a key point detection model to determine the skeleton points of an athlete in the first action image, including step S1210. In step S1210, the first target image is input into the key point detection model to obtain the skeleton points of the athlete. The key point detection model detects the first target image to determine the skeleton points of the athlete.
[0101] The above technical solution uses a target detection model to determine the predicted position of the athlete in a first action image. This first action image is then processed to obtain a first target image. This first target image is then detected using a key point detection model to determine the athlete's skeleton points. This reduces background image interference compared to the first action image, improving the accuracy of the key point detection model in detecting the athlete's skeleton points and enhancing the user experience.
[0102] Exemplarily, the target detection model is obtained by performing a joint training operation using multiple groups of third action images and position annotation information of the athletes in each third action image, wherein each group of third action images includes third action images of the same target object acquired synchronously from multiple angles.
[0103] The third action image can be the same as the above-mentioned second action image, and both are obtained by synchronously shooting the athlete using a camera device at multiple angles. The position of the athlete in the third action image can be marked to obtain the athlete's position annotation information. The athlete's position can be marked in the form of a human body bounding box. Similarly, the target detection model can use the third action images of multiple angles obtained synchronously and the athlete's position annotation information in each of the third action images to perform a joint training operation. The target detection model can be jointly trained according to a loss function calculation method similar to the above-mentioned first loss function value, second loss function value and / or third loss function value, and the specific details are not repeated here. The trained target detection model can be used as the target detection model for determining the position prediction information of the athlete in the first action image in the above-mentioned step S1400.
[0104] The above technical solution trains the target detection model by synchronously acquiring third action images from multiple angles and the corresponding position annotation information of the athletes. This can improve the target detection model's ability to detect athletes at different angles and postures, and improve the accuracy of the target detection model.
[0105] Illustratively, according to another aspect of the present invention, a device for identifying skeleton points of an athlete is provided. Figure 7 FIG. 7 shows a schematic block diagram of a skeleton point recognition device 700 for an athlete according to an embodiment of the present application. Figure 7As shown, the athlete's skeleton point recognition device 700 includes a first acquisition module 710 and a first determination module 720 .
[0106] The first acquisition module 710 is configured to acquire a first motion image of an athlete. The first determination module 720 is configured to determine the skeletal points of the athlete in the first motion image using a key point detection model based on the first motion image. The key point detection model is obtained by performing a joint training operation using multiple sets of second motion images and positional annotation information of the skeletal points in each of the second motion images, where each set of second motion images includes synchronously acquired second motion images of the same target object from multiple angles.
[0107] Exemplarily, before determining the skeletal point information of the athlete in the first action image using a key point detection model based on the first action image, the athlete's skeletal point identification device 700 also includes a second determination module and a processing module. The first determination module includes a first determination submodule. The second determination module is used to determine the position prediction information of the athlete in the first action image based on the first action image using a target detection model. The processing module is used to process the first action image based on the position prediction information of the athlete in the first action image to obtain a first target image corresponding to the first action image, wherein the first target image includes the area where the athlete is located. The first determination submodule is used to input the first target image into the key point detection model to obtain the skeletal point information of the athlete.
[0108] Exemplarily, the target detection model is obtained by performing a joint training operation using multiple groups of third action images and position annotation information of the athletes in each of the third action images, wherein each group of third action images includes third action images of the same target object acquired synchronously from multiple angles.
[0109] Exemplarily, the apparatus 700 for identifying skeletal points of an athlete includes a joint training module. The joint training module includes a first conversion submodule, a second determination submodule, and a first training submodule. The first conversion submodule is used to convert the position annotation information of the skeletal points in a second action image at another angle, which is acquired synchronously with the second action image at the first angle, into position information at the first angle, as auxiliary supervision data for the second action image at the first angle, wherein the first angle is one of a plurality of angles. The second determination submodule is used to input the second action image at the first angle into a key point detection model to obtain position prediction information of the skeletal points of the athlete in the second action image at the first angle. The first training submodule is used to train the key point detection model according to a first loss function value, wherein the first loss function value is calculated based on at least the position prediction information, the first main supervision data, and the auxiliary supervision data of the second action image at the first angle, wherein the first main supervision data includes the position annotation information of the skeletal points of the athlete in the second action image at the first angle.
[0110] Exemplarily, for each second angle among the multiple angles that is different from the first angle, the joint training module also includes a second conversion submodule and a third determination submodule. The second conversion submodule is used to convert the position annotation information of the skeleton points in the second action image at other angles other than the second angle, which is obtained synchronously with the second action image at the second angle, into position information at the second angle, as auxiliary supervision data for the second action image at the second angle. The third determination submodule is used to input the second action image at the second angle into a key point detection model to obtain the position prediction information of the skeleton points of the athlete in the second action image at the second angle. The first loss function value in the first training submodule is also calculated based on the position prediction information of the skeleton points of the athlete in the second action image at the second angle, the second main supervision data and the auxiliary supervision data of the second action image at the second angle, and the second main supervision data includes the position annotation information of the skeleton points of the athlete in the second action image at the second angle.
[0111] Exemplarily, the joint training module includes a fourth determination submodule, a fifth determination submodule, and a second training submodule. The fourth determination submodule is used to input at least the first image in the second action image into the key point detection model to obtain position prediction information of the skeletal points of the athlete in the first image. The fifth determination submodule is used to determine the polar line corresponding to each skeletal point in the first image based on the basic matrix corresponding to the second image in the first image and the second action image, wherein the second image is an image acquired synchronously with the first image at a different angle, and the basic matrix represents the geometric relationship between the first image and the second image. The second training submodule is used to train the key point detection model according to the second loss function value, wherein the second loss function value is calculated based on the distance from the position prediction information of each skeletal point in the first image to its corresponding polar line.
[0112] Exemplarily, the joint training module includes a sixth determination submodule and a third training submodule. The sixth determination submodule is used to input at least one second motion image into the key point detection model to obtain position prediction information of the skeleton points of the athlete in the second motion image. The third training submodule is used to train the key point detection model according to a third loss function value, wherein the third loss function value is based on the position prediction information of the skeleton points in the second motion image, the position annotation information of the skeleton points in the second motion image and other second motion images, and the skeleton point weight calculation, wherein among all the skeleton points in the second motion image, there are at least two skeleton points with different skeleton point weights.
[0113] In an exemplary embodiment, the athlete's skeleton point recognition apparatus 700 further includes a classification and evaluation module configured to classify and / or evaluate the athlete's motion in the first motion image based on the athlete's skeleton points in the first motion image.
[0114] Illustratively, according to yet another aspect of the present invention, an electronic device is further provided. Figure 8 A schematic block diagram of an electronic device 800 according to an embodiment of the present application is shown. The electronic device 800 includes a processor 810 and a memory 820. The memory 820 stores computer program instructions, which are used by the processor 810 to execute the above-mentioned method for identifying the skeleton points of an athlete.
[0115] Illustratively, according to another aspect of the present invention, a storage medium is provided, on which program instructions are stored, which are used to execute the above-mentioned athlete's skeletal point recognition method when running. The storage medium may include, for example, an erasable programmable read-only memory (EPROM), a portable CD-ROM, a USB memory, or any combination of the above storage media. The storage medium may be any combination of one or more computer-readable storage media.
[0116] Illustratively, according to yet another aspect of the present invention, a computer program product is provided, comprising computer program instructions, which are used to execute the above-mentioned method for identifying skeletal points of an athlete when running.
[0117] A person skilled in the art can understand the specific implementation schemes and beneficial effects of the above-mentioned athlete's skeleton point recognition device, electronic device, storage medium and computer program product by reading the above-mentioned description of the athlete's skeleton point recognition method. For the sake of brevity, they will not be repeated here.
[0118] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0119] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0120] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.
[0121] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0122] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0123] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0124] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0125] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the skeletal point identification device for a player according to an embodiment of the present application. The present application can also be implemented as a device program (e.g., a computer program and a computer program product) for executing a part or all of the method described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0126] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0127] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for identifying skeleton points of an athlete, characterized in that: include: Acquiring a first motion image of the athlete; Based on the first action image, a key point detection model is used to determine the skeleton point information of the athlete in the first action image, wherein the key point detection model is obtained by performing a joint training operation using multiple groups of second action images and the position annotation information of the skeleton points in each of the second action images, and each group of second action images includes second action images of the same target object acquired synchronously from multiple angles.
2. The method for identifying skeleton points of an athlete according to claim 1, wherein: Before determining the skeleton point information of the athlete in the first action image using a key point detection model based on the first action image, the athlete skeleton point recognition method further includes: Based on the first action image, using a target detection model, determining predicted position information of the athlete in the first action image; processing the first action image according to the predicted position information of the athlete in the first action image to obtain a first target image corresponding to the first action image, wherein the first target image includes an area where the athlete is located; The determining, based on the first action image and using a key point detection model, the skeleton points of the athlete in the first action image includes: The first target image is input into the key point detection model to obtain the skeleton point information of the athlete.
3. The method for identifying skeleton points of an athlete according to claim 2, wherein: The target detection model is obtained by performing a joint training operation using multiple groups of third action images and position annotation information of the athletes in each of the third action images, wherein each group of third action images includes third action images of the same target object acquired synchronously at multiple angles.
4. The method for identifying skeleton points of an athlete according to claim 1, wherein: The joint training operation includes: For a second action image at another angle that is acquired synchronously with the second action image at the first angle, converting position annotation information of skeletal points in the second action image into position information at the first angle as auxiliary supervision data for the second action image at the first angle, wherein the first angle is one of the multiple angles; Inputting the second motion image from the first angle into the key point detection model to obtain position prediction information of the skeleton points of the athlete in the second motion image from the first angle; The key point detection model is trained according to a first loss function value, wherein the first loss function value is calculated based on at least the position prediction information, first main supervision data and auxiliary supervision data of the second action image at the first angle, and the first main supervision data includes position annotation information of the skeleton points of the athlete in the second action image at the first angle.
5. The method for identifying skeleton points of an athlete according to claim 4, wherein: The joint training operation further includes: For each second angle of the plurality of angles that is different from the first angle, Converting position annotation information of skeletal points in second motion images at angles other than the second angle, which are acquired synchronously with the second motion image at the second angle, into position information at the second angle as auxiliary supervision data for the second motion image at the second angle; Inputting the second action image at the second angle into the key point detection model to obtain position prediction information of the skeleton points of the athlete in the second action image at the second angle; In which, the first loss function value is also calculated based on the position prediction information of the skeleton points of the athlete in the second action image at the second angle, the second main supervision data and the auxiliary supervision data of the second action image at the second angle, and the second main supervision data includes the position annotation information of the skeleton points of the athlete in the second action image at the second angle.
6. The method for identifying skeleton points of an athlete according to claim 1, wherein: The joint training operation includes: Inputting at least a first image in the second action image into the key point detection model to obtain position prediction information of skeletal points of the athlete in the first image; determining an epipolar line corresponding to each skeletal point in the first image based on a fundamental matrix corresponding to a second image in the first image and the second action image, wherein the second image is an image acquired synchronously with the first image at a different angle, and the fundamental matrix represents a geometric relationship between the first image and the second image; The key point detection model is trained according to a second loss function value, wherein the second loss function value is calculated based on the distance between the position prediction information of each skeleton point in the first image and its corresponding epipolar line.
7. The method for identifying skeleton points of an athlete according to claim 1, wherein: The joint training operation includes: Inputting at least one second motion image into the key point detection model to obtain position prediction information of the skeleton points of the athlete in the second motion image; The key point detection model is trained according to a third loss function value, wherein the third loss function value is based on position prediction information of the skeleton points in the second action image, position annotation information of the skeleton points in the second action image and other second action images, and skeleton point weight calculation, wherein among all the skeleton points in the second action image, there are at least two skeleton points with different skeleton point weights.
8. An electronic device comprising: A processor and a memory, characterized in that The memory stores computer program instructions, which are used by the processor to execute the method for identifying skeletal points of an athlete according to any one of claims 1 to 7 when the processor is running the computer program instructions.
9. A storage medium having program instructions stored thereon, characterized in that: The program instructions are used to execute the athlete's skeleton point recognition method according to any one of claims 1 to 7 when running.
10. A computer program product comprising computer program instructions, characterized in that The computer program instructions are used to execute the athlete's skeleton point recognition method according to any one of claims 1 to 7 when running.
Citation Information
Patent Citations
Interactive behavior recognizing method, device, computer equipment and storage medium
CA3154025A1
Multi-angle tumble high-risk identification method and system based on skeleton key points
CN113496216A
Action recognition method and device based on skeleton point distance, equipment and storage medium
CN114724241A
Multi-person motion detection method and device based on human skeleton point detection
CN115690895A