Action evaluation method, device, equipment, medium and program product
By integrating two-dimensional image coordinates with depth information, the method improves dance evaluation accuracy by comparing key points in three-dimensional space, addressing the inaccuracies of two-dimensional scoring methods.
Patent Information
- Application Number
- CN202510367759.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, dance action scoring methods based on plane coordinate information are prone to misjudgment in complex movements, resulting in low scoring accuracy and it is difficult to fully capture the three-dimensional spatial characteristics of the movement.
Movement evaluation is performed based on the depth information of the limb key points, and the three-dimensional information of the limb key points of the first subject and the reference action is obtained to determine the degree of action matching, and the deep learning model is used to analyze the differences between the limb key points in the three-dimensional space.
It improves the accuracy and comprehensiveness of movement evaluation, reduces the error caused by inconsistent viewing angle and longitudinal distance, and can more accurately detect the movement trajectory of key points of the limb.
Smart Images

Figure CN120318902A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular, to a method, device, equipment, medium, and program product for evaluating actions. Background Art
[0002] In the fields of dance teaching and games, introducing a dance scoring mechanism to evaluate the standard degree of users' dance movements can help users understand their own dance levels, improve dance skills, and enhance the learning or gaming experience of users.
[0003] In related technologies, an image containing a user's dance movement is collected through a camera, and a deep learning model (such as OpenPose, a human pose estimation model) is used to evaluate the user's action. The deep learning model detects the image to obtain the planar coordinates of the human key points, analyzes the distance and angle changes between the points, and compares and analyzes with the standard action to score the accuracy of the user's action.
[0004] However, the above scoring strategy only relies on planar coordinate information for simple calculations. For complex dance movements, if the distance between key points is close and easy to confuse, there may be misjudgments, resulting in a low accuracy of the scoring result. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, equipment, medium, and program product for evaluating actions, which can evaluate limb actions by combining the depth information of limb key points and improve the accuracy of the action evaluation result. The technical solution is as follows:
[0006] On the one hand, a method for evaluating an action is provided, and the method includes:
[0007] Obtain a first image, where the first image includes a first subject performing a target action;
[0008] Based on the first image, determine limb key point data for characterizing the target action, where the limb key point data includes a first image coordinate and first depth information corresponding to a first limb key point of the first subject;
[0009] Obtain reference key point data for characterizing a reference action corresponding to the target action, where the reference key point data includes a second image coordinate and second depth information corresponding to a second limb key point of a second subject, and the second subject is the execution subject of the reference action;
[0010] Based on the limb key point data and the reference key point data, determine an evaluation result for the target action.
[0011] On the other hand, an apparatus for evaluating an action is provided, and the apparatus includes:
[0012] An image acquisition module, configured to acquire a first image, where the first image includes a first subject performing a target action;
[0013] A data acquisition module, configured to determine limb key point data for characterizing the target action based on the first image, where the limb key point data includes a first image coordinate and a first depth information corresponding to a first limb key point of the first subject;
[0014] The data acquisition module is further configured to acquire reference key point data for characterizing a reference action corresponding to the target action, where the reference key point data includes a second image coordinate and a second depth information corresponding to a second limb key point of a second subject, and the second subject is an execution subject of the reference action;
[0015] An evaluation module, configured to determine an evaluation result of the target action based on the limb key point data and the reference key point data.
[0016] In an optional embodiment, the evaluation module is further configured to perform a similarity analysis on the limb key point data and the reference key point data, and obtain a similarity score as the evaluation result of the target action.
[0017] In an optional embodiment, the evaluation module is further configured to obtain a first similarity between the first image coordinate and the second image coordinate; obtain a second similarity between the first depth information and the second depth information; and obtain the similarity score based on a weighted sum between the first similarity and the second similarity.
[0018] In an optional embodiment, the evaluation module is further configured to obtain a first coordinate similarity between a first coordinate of the first image coordinate in a first axis direction and a first coordinate of the second image coordinate in the first axis direction; and obtain a second coordinate similarity between a second coordinate of the first image coordinate in a second axis direction and a second coordinate of the second image coordinate in the second axis direction; and obtain the first similarity based on a weighted sum between the first coordinate similarity and the second coordinate similarity.
[0019] In an optional embodiment, the limb key points include at least two types of key points representing different limb parts; the i-th type of key point in the limb key points corresponds to the i-th limb part of the first subject and the second subject, and i is a positive integer;
[0020] The evaluation module is further configured to, for the i-th type of key points, obtain an evaluation sub-result for the i-th limb part based on the difference between the i-th sub-data in the limb key point data and the i-th reference sub-data in the reference key point data; and obtain the evaluation result for the target action based on the evaluation sub-results corresponding to the at least two types of key points respectively.
[0021] In an optional embodiment, the apparatus further includes:
[0022] A display module, configured to display a program interface of a first program, where the first program is used to provide an action evaluation function for the first entity; and in response to receiving an action evaluation start operation, play a reference video at a first moment, where the reference video is used to show a picture of the second entity performing the reference action, and the reference video includes at least two frames of images; where the first image is an image collected by the terminal at a second moment, and the first moment is a moment before the second moment.
[0023] In an optional embodiment, the data acquisition module is further configured to obtain frame sequence data based on the reference video, where the frame sequence data includes the at least two frames of images in the reference video and time stamps respectively corresponding to the at least two frames of images, and the time stamps are used to indicate the progress of each frame of image in the reference video; and obtain a key point data set of the second entity based on the frame sequence data, where the key point data set includes the reference key point data, and the key point data set includes the image coordinates and depth information corresponding to the second limb key points when the second entity performs the reference action in each frame of image in the reference video.
[0024] In an optional embodiment, the data acquisition module is further configured to perform key point analysis on each frame of image in the frame sequence data respectively to obtain key point data corresponding to the at least two frames of images respectively; and integrate the key point data corresponding to the at least two frames of images respectively based on the time stamps corresponding to the at least two frames of images respectively to obtain the key point data set.
[0025] In an optional embodiment, the data acquisition module is further configured to obtain a time difference between the second moment and the first moment, where the time difference is used to indicate the playing progress of the reference video at the second moment; and obtain the reference key point data from the key point data set based on the time difference and the time stamps included in the key point data set, where the reference key point data corresponds to a second image in the at least two frames of images, and the progress of the second image in the reference video matches the playing progress of the reference video at the second moment.
[0026] In an alternative embodiment, the evaluation module is further configured to convert the evaluation result into an evaluation score, and the evaluation score is used to describe the evaluation result in numerical form;
[0027] The display module is further configured to display the evaluation score in a first prominent form in the result display area of the first program, and there is a corresponding relationship between the type of the first prominent form and the value of the evaluation score.
[0028] In an alternative embodiment, the device further includes:
[0029] A mapping module, configured to, in response to receiving an action mapping operation, map the action of the second subject in the reference video to the first subject in the first image based on the first limb key points and the second limb key points, where the first limb key points and the second limb key points meet a preset mapping coincidence requirement.
[0030] On the other hand, a computer device is provided, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the action evaluation method according to any one of the above embodiments of the present application.
[0031] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the action evaluation method according to any one of the above embodiments of the present application.
[0032] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the action evaluation method according to any one of the above embodiments.
[0033] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0034] When evaluating the target action performed by the first subject, the reference action performed by the second subject is used as the evaluation criterion. By obtaining the three-dimensional information of the limb key points of the first subject and the second subject in the image respectively, the matching degree between the reference action and the target action is determined based on the difference between the three-dimensional information, and the evaluation result of the target action is obtained. The three-dimensional information includes the planar coordinates and depth information of the limb key points, which makes the positioning of the limb key points in the three-dimensional space more accurate. Compared with the method of performing action matching and evaluation only based on planar coordinates, it can reduce the influence of action deformation caused by perspective problems and inconsistent longitudinal distances in the planar image, detect the movement trajectory of the limb key points more accurately, and improve the accuracy and comprehensiveness of the action evaluation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0036] Figure 1 is a schematic diagram of an action evaluation system provided by an exemplary embodiment of the present application;
[0037] Figure 2 is a flowchart of an action evaluation method provided by an exemplary embodiment of the present application;
[0038] Figure 3 is a schematic diagram of dividing limb key points provided by an exemplary embodiment of the present application;
[0039] Figure 4 is a flowchart of an action evaluation method provided by an exemplary embodiment of the present application;
[0040] Figure 5 is a schematic diagram of the display mode of the evaluation result provided by an exemplary embodiment of the present application;
[0041] Figure 6 is a flowchart of an action following method provided by an exemplary embodiment of the present application;
[0042] Figure 7 is a structural block diagram of an action evaluation device provided by an exemplary embodiment of the present application;
[0043] Figure 8 is a structural block diagram of an action evaluation device provided by another exemplary embodiment of the present application;
[0044] Figure 9It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0046] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0047] It should be noted that the information and data involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0048] First, a brief introduction to the nouns involved in the embodiments of the present application:
[0049] Artificial Intelligence (AI): It is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0050] Machine Learning (ML): It is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0051] AI model (artificial intelligence model): It refers to a computer program constructed through technologies such as machine learning and deep learning that can analyze, infer, and predict data. Based on a large amount of training data and algorithms, it can automatically learn from the input data, extract features, and make decisions and predictions based on the learned knowledge.
[0052] In this application, an image containing a first subject and a reference video containing a second subject are input into a pre-trained AI model, and the AI model outputs the limb key point data of the first subject and the reference key point data of the second subject.
[0053] Among them, the AI model can directly output key point data that simultaneously includes depth information and planar coordinates; it can also only output planar coordinates, obtain depth information when the depth camera captures images, and integrate the depth information and planar coordinates to obtain key point data. This application does not limit this.
[0054] In the fields of dance teaching, interactive entertainment, and game applications, introducing a dance movement scoring mechanism can effectively evaluate the standard degree of the user's dance movements, thereby providing accurate feedback to the user, helping them deeply understand their own dance level, continuously improve their dance skills, and significantly enhance the user's learning experience or game immersion.
[0055] In related technologies, image data containing the user's dance movements is mainly collected through a camera, and a deep learning model is used to evaluate the user's movements. For example, OpenPose is a real-time multi-person two-dimensional body pose estimation model based on deep learning, which can be used for real-time detection of key points of the body, feet, hands, and face.
[0056] Specifically, the deep learning model detects the image, extracts the two-dimensional planar coordinate information of the human body limb nodes, further analyzes the distance and angle changes between the nodes, and finally quantitatively scores the accuracy of the user's movements through comparative analysis with the standard movements.
[0057] However, the existing scoring strategies mainly rely on two-dimensional planar coordinate information for simple geometric calculations, and this method has limitations. For complex dance movements, especially when the distance between limb nodes is close or the movements overlap, it is easy to have the situation of limb point confusion or misjudgment, resulting in a decrease in the accuracy of the scoring results. In addition, it is difficult to comprehensively capture the three-dimensional spatial characteristics of the movements only relying on planar coordinate information, and it is impossible to fully reflect details such as the fluency of dance movements, limiting the scoring effect. Therefore, there is an urgent need for a more accurate and comprehensive dance movement scoring method to improve the user experience and the reliability of the scoring system.
[0058] Based on the above problems, this application provides a method for evaluating actions, which can comprehensively evaluate the matching degree between the actions performed by the subject and the reference actions based on the two-dimensional planar coordinates and depth information of the limb key points in the image. Adding depth information can reduce the errors caused by different depths of the limb key points and improve the accuracy of the evaluation results of the actions.
[0059] Secondly, the implementation environment involved in the embodiments of this application is described. Schematically, please refer toFigure 1 , Figure 1 is a schematic diagram of an action evaluation system. The system involves a terminal 100, and a first program runs in the terminal 100. The first program can provide an action evaluation function.
[0060] Obtain a reference video. The reference video includes images of a second subject performing a reference action. The reference video is the standard for action evaluation and contains multiple video frames.
[0061] Extract frame sequence data 110 from the reference video. The frame sequence data 110 is a sequence composed of video frames. The frame sequence data 110 contains each video frame and its corresponding timestamp, and the timestamp is used to indicate the progress of the video frame in the reference video.
[0062] Input the frame sequence data 110 into the AI model 120. The AI model 120 performs key point analysis on the frame sequence data 110, extracts information about the limb key points in each frame image, and obtains a key point data set. Among them, the key point data set includes the image coordinates and depth information corresponding to the second limb key points of the second subject in each video frame, and the timestamp corresponding to each video frame.
[0063] Store the key point data set in the local database 130. When a user (hereinafter referred to as the first subject) starts the first program through the terminal 100, starts a live broadcast and enables the action evaluation function, the program interface of the first program will be blurred and the reference video will be played, providing action prompts for the first subject and determining the action evaluation result for the first subject as a benchmark. That is, the program interface simultaneously contains the image of the first subject's action and the blurred reference video.
[0064] Exemplarily, play the reference video at the first moment.
[0065] When the first subject performs a target action based on the reference action of the second subject, the terminal 100 will collect the current frame (the live broadcast image at the current moment) in real time through the camera component, and input the current frame into the AI model 120 to obtain information about the limb key points in the current frame.
[0066] For example, when the current moment is the second moment, the live broadcast image collected is the first image, and the first image contains the image of the first subject performing the target action. Input the first image into the AI model 120, and output the limb key point data of the first subject.
[0067] Among them, the data format of the limb key point data corresponds to that of each video frame in the key point data set. The limb key point data of the first subject includes the first image coordinates and the first depth information corresponding to the first limb key points of the first subject.
[0068] For example, the number of key limb points of the first body is 33, including 11 head limb points, 12 hand limb points, and 10 leg limb points. The first image coordinates include the two-dimensional coordinates of each key limb point in the first image, and the first depth information includes the depth of each key limb point in the first image.
[0069] Based on the timestamps corresponding to the first moment and the second moment, the video frame (second image) corresponding to the first image is matched, and the reference key point data of the second image is obtained from the key point data set.
[0070] The similarity analysis is performed on the reference key point data and the key limb point data to obtain a similarity score of 140, which is used as the evaluation result of the target action in the first image. The evaluation result is used to express the matching degree between the target action in the first image and the reference action in the second image.
[0071] Among them, in the similarity analysis process, the similarity calculations need to be performed on the head limb points, hand limb points, and leg limb points respectively to obtain the similarities corresponding to the head action, hand action, and leg action. Then, based on the three similarities, a weighted operation is performed to obtain the similarity score of 140, which is used as the evaluation result of the overall action.
[0072] In some embodiments, the system also involves a server, in which an AI model 120 is deployed, and there is a communication connection between the server and the terminal 100.
[0073] The above terminal 100 can be various forms of terminal devices such as mobile phones, tablet computers, desktop computers, portable laptops, smart TVs, vehicle-mounted terminals, and smart home devices, and the embodiments of the present application do not limit this.
[0074] It should be noted that the above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0075] In some embodiments, the above server can also be implemented as a node in a blockchain system.
[0076] Combined with the above noun introduction and application scenarios, the method for evaluating the action provided by the present application is described. This method can be executed by the server or the terminal, or jointly executed by the server and the terminal. In the embodiments of the present application, an example in which this method is executed by the terminal is used for description, as Figure 2 shownFigure 2 It is a flowchart of a method for evaluating actions provided by an exemplary embodiment of the present application. The method includes the following steps.
[0077] Step 210, obtain a first image.
[0078] The first image includes a first subject performing a target action.
[0079] The first subject refers to an individual with a complete limb structure, capable of moving or performing certain actions. The types of the first subject include but are not limited to: (1) humans; (2) animals; (3) robots with limb forms, etc.
[0080] Exemplarily, in this embodiment, the first subject is taken as an example of a human for illustration.
[0081] The types of the first image include but are not limited to the following several types: (1) an image collected in real time by a camera component of the terminal at the current moment, used to show the current activity state of the first subject; (2) an image collected or received by the terminal at a historical moment (i.e., a moment before the current moment), used to show the activity state of the first subject at the historical moment.
[0082] Among them, the first image meets the preset limb key point display requirement, and the limb key point display requirement means that the number of limb key points of the subject shown in the image reaches a preset key point number threshold.
[0083] The limb key point is a key point obtained by marking a specified position in the limb part of the subject.
[0084] Step 220, determine limb key point data for characterizing the target action based on the first image.
[0085] Among them, the limb key point data includes the first image coordinates and the first depth information corresponding to the first limb key points of the first subject.
[0086] The first image coordinates refer to the coordinates of the first limb key point in the coordinate system when the first image is a two-dimensional coordinate system, and the first depth information refers to the depth of the first limb key point from the camera that collects the first image.
[0087] Among them, the number of the first limb key points recognized in the first image is at least two, the first image coordinates include the coordinates of each first limb key point, and the first depth information includes the depth of each first limb key point.
[0088] That is, the limb key point data is a four-dimensional tensor Q ∈ R N×U×V×D , where R is a real number, N is the number of limb key points, U and V represent the position coordinates of the limb key points in the two-dimensional space (the first image), and D represents the depth of the limb key points.
[0089] Exemplarily, perform key point analysis on the first image, input the first image into a pre-trained AI model, and extract the limb key point data of the first subject.
[0090] The pre-trained AI model can identify the limb key points of the subject contained in the image and analyze the two-dimensional coordinates and depth information of the limb key points in the image.
[0091] In some embodiments, it is also possible to use a pre-trained AI model to identify limb key points and analyze the two-dimensional coordinates of the limb key points in the image. When collecting the first image, use a depth camera (a device capable of obtaining the depth information of objects in the scene), and the depth information of each limb key point can be directly obtained. The manner of obtaining the limb key point data is not limited in this embodiment.
[0092] Step 230, obtain reference key point data for characterizing the reference action corresponding to the target action.
[0093] Among them, the reference key point data includes the second image coordinates and the second depth information corresponding to the second limb key points of the second subject, and the second subject is the execution subject of the reference action.
[0094] The limb key point data corresponding to the first subject and the reference key point data corresponding to the second subject respectively represent the actual execution situations of the first subject and the second subject when performing the same action.
[0095] The manner of obtaining the reference key point data and the data format are the same as those of the limb key point data of the first subject. The reference key point data is a reference standard for evaluating whether the target action performed by the first subject conforms to the reference action.
[0096] Step 240, determine the evaluation result of the target action based on the limb key point data and the reference key point data.
[0097] Among them, the evaluation result is used to express the matching degree between the target action and the reference action.
[0098] Optionally, perform similarity analysis on the limb key point data and the reference key point data to obtain a similarity score as the evaluation result of the target action.
[0099] 1. Obtain the first similarity between the first image coordinates and the second image coordinates.
[0100] Among them, both the first image coordinates and the second image coordinates are two-dimensional coordinates, and there are components in both the first axis direction and the second axis direction. The first axis direction is different from the second axis direction. In this embodiment, the first axis direction and the second axis direction are two mutually perpendicular directions.
[0101] Exemplarily, the first image is rectangular. A specified point in the first image is determined as the origin, the horizontal direction corresponding to the bottom side of the first image is determined as the first axis direction, and the vertical direction corresponding to the side of the first image is determined as the second axis direction. An image coordinate system is established to obtain the first image coordinates.
[0102] The second image coordinates are determined in the same way as above, and a coordinate system is established based on the second image corresponding to the reference key point data.
[0103] In some embodiments, in order to eliminate the image coordinate error caused by the different sizes of the first image and the second image, and the depth information error caused by the different distances between the subject and the camera when collecting images, when obtaining the limb key point data, the first image is preprocessed so that the size and depth information of the first image match those of the second image.
[0104] For example, identify the limb key points in the first image and the second image. At least two relatively stable limb key points (such as the head limb points) in the subject are used as reference points. Taking the relative positions between the reference points of the second subject as the target, the reference points of the second subject are respectively labeled with numbers 1 to k, where k is a positive integer. For the first image, mapping points corresponding one by one to the reference points of the second subject are obtained from the first image and are respectively labeled with numbers 1 to k. That is, there is a corresponding relationship between the reference point with number 1 of the second subject and the mapping point with number 1.
[0105] The reference point and the mapping point with the same number are called a reference point pair. For example, the reference point with number k of the second subject and the mapping point with number k are called reference point pair k.
[0106] The first image and the second image are overlapped so that the reference points and the mapping points are in the same two-dimensional plane, and the distances between each reference point pair in this two-dimensional plane are calculated respectively. For reference point pair k, the coordinate k1 of reference point k in the two-dimensional plane and the coordinate k2 of mapping point k in the two-dimensional plane are obtained, and the distance between k1 and k2 is calculated based on the distance formula between two points. The mean value between the distances of each pair of reference point pairs is calculated.
[0107] The size of the first image is adjusted in real time. After each adjustment of the size of the first image, the mean value between the distances of each pair of reference point pairs is updated based on the above steps to obtain the updated distance mean value until the distance mean value is less than or equal to a preset distance threshold, and then the adjustment of the size of the first image is stopped. At this time, the relative positions between the reference points of the first subject can be made to match the relative positions between the reference points of the second subject.
[0108] Analyze the first image to obtain the angle parameters and distance parameters recorded respectively when the first image and the second image are captured by different cameras. Use image transformation algorithms, such as perspective transformation, scaling transformation, etc., to correct the first image so that the corrected first image presents the visual effect obtained by capturing the second image from the same perspective.
[0109] Optionally, obtain the first coordinate similarity between the first coordinate of the first image coordinate in the first axis direction and the first coordinate of the second image coordinate in the first axis direction. And, obtain the second coordinate similarity between the second coordinate of the first image coordinate in the second axis direction and the second coordinate of the second image coordinate in the second axis direction.
[0110] The methods for calculating the first coordinate similarity and the second coordinate similarity are the same.
[0111] Based on the weighted sum between the first coordinate similarity and the second coordinate similarity, obtain the first similarity.
[0112] Among them, when performing the weighted operation, the first weight corresponding to the first coordinate similarity and the second weight corresponding to the second coordinate similarity are preset weights.
[0113] For example, in this embodiment, both the first weight and the second weight are 0.5.
[0114] 2. Obtain the second similarity between the first depth information and the second depth information.
[0115] Among them, the method for calculating the second similarity is the same as the method for calculating the first coordinate similarity.
[0116] 3. Obtain the similarity score based on the weighted sum between the first similarity and the second similarity.
[0117] Among them, when performing the weighted operation, the third weight corresponding to the first similarity and the fourth weight corresponding to the second similarity are preset weights.
[0118] For example, in this embodiment, both the third weight and the fourth weight are 0.5.
[0119] Optionally, the limb key points include at least two types of key points representing different limb parts. The i-th type of key point in the limb key points corresponds to the i-th limb part of the first subject and the second subject, where i is a positive integer.
[0120] For the i-th type of key point, based on the difference between the i-th sub-data in the limb key point data and the i-th reference sub-data in the reference key point data, obtain the evaluation sub-result for the i-th limb part.
[0121] Wherein, the i-th sub-data is the data corresponding to the i-th type of key points in the limb key point data, and the i-th reference sub-data is the data corresponding to the i-th type of key points in the reference key point data.
[0122] For example, if the i-th type of key points corresponds to the heads of the first main body and the second main body, then the i-th sub-data refers to the data related to the head of the first main body in the limb key point data, and the i-th reference sub-data refers to the data related to the head of the second main body in the reference key point data.
[0123] Exemplarily, the method for obtaining the evaluation sub-result of the i-th limb part refers to the method for calculating the similarity score described above. That is, for each limb part, the sub-data corresponding to the limb part is respectively obtained from the limb key point data and the reference key point data, and the similarity analysis is performed on the sub-data, and the obtained analysis result is used as the evaluation sub-result of the i-th limb part.
[0124] Based on the evaluation sub-results corresponding to at least two types of key points respectively, the evaluation result of the target action is obtained.
[0125] Exemplarily, the evaluation sub-results of each type of key points are integrated to obtain the complete evaluation result when the first main body performs the target action. For example, if the evaluation sub-results corresponding to at least two types of key points respectively are all expressed as similarity values, then the weighted sum of the similarity values is calculated to obtain the evaluation result of the target action.
[0126] Exemplarily, the limb key points include head key points, hand key points and leg key points, and the number of head key points.
[0127] Schematically, as Figure 3 shown, Figure 3 is a schematic diagram for dividing limb key points.
[0128] Taking the first main body and the second main body as humans for illustration, each main body contains 33 limb key points, which are divided into 310 head limb points (serial numbers from 0 to 10), 320 hand limb points (serial numbers from 11 to 22), and 330 leg limb points (serial numbers from 23 to 32).
[0129] When analyzing the matching degree between the target action performed by the first main body and the reference action performed by the second main body to obtain the evaluation result, first, the similarity is calculated for the three types of limb points respectively as the evaluation sub-results for different limb parts, and then the final complete evaluation result of the target action is obtained based on the weighted operation of the similarities corresponding to the three types of limb points respectively.
[0130] Exemplarily, taking the calculation of the similarity of the head limb points as an example for illustration, the method for calculating the similarity of the hand limb points and the leg limb points is the same.
[0131] Calculate the similarity C of the head limb points based on the following Formulas 1 to 5 head .
[0132] Formula 1:
[0133]
[0134] where refers to the similarity of the head limb points of the first subject and the head limb points of the second subject in the first axis direction.
[0135] is any one of them, u Q is the vector formed by the coordinates of the head limb points in the first axis direction in the limb key point data. For example, if the image coordinates corresponding to the first head limb point in the limb key point data are (x1, y1), then is the component x1 of the first head limb point in the first axis direction, and x1 and y1 are real numbers.
[0136] refers to any one of them, u P is the vector formed by the coordinates of the head limb points in the first axis direction in the reference key point data. For example, if the image coordinates corresponding to the first head limb point in the reference key point data are (x, y), then is the component x of the first head limb point in the first axis direction, and x and y are real numbers.
[0137] Formula 2:
[0138]
[0139] where refers to the similarity of the head limb points of the first subject and the head limb points of the second subject in the second axis direction.
[0140] is any one of them, v Q is the vector formed by the coordinates of the head limb points in the second axis direction in the limb key point data. For example, if the image coordinates corresponding to the first head limb point in the limb key point data are (x1, y1), then is the component y1 of the first head limb point in the second axis direction, and x1 and y1 are real numbers.
[0141] refers to any one of them, is a vector composed of the coordinates of the head and limb points in the reference key point data in the second axis direction. For example, if the image coordinates corresponding to the first head and limb point in the reference key point data are (x, y), then is the component y of the first head and limb point in the second axis direction, and x and y are real numbers.
[0142] Formula Three:
[0143]
[0144] C head1 refers to the first similarity between the image coordinates of the head and limb points of the first subject and the head and limb points of the second subject.
[0145] Perform a weighted operation on the similarities of the head and limb points of the first subject and the head and limb points of the second subject in the first axis direction and the second axis direction respectively to obtain the first similarity between the image coordinates of the head and limb points.
[0146] Exemplarily, the weights assigned to the similarities of the head and limb points in the two axis directions in Formula Three are both 0.5. In some embodiments, the value of the weight can be any real number.
[0147] Formula Four:
[0148]
[0149] Among them, refers to the second similarity between the depth information of the head and limb points of the first subject and the head and limb points of the second subject.
[0150] is any one of them, d Q is a vector composed of the depths of the head and limb points in the limb key point data.
[0151] refers to any one of them, d P is a vector composed of the depths of the head and limb points in the reference key point data.
[0152] Formula Five:
[0153]
[0154] Among them, C head is the similarity between the head and limb points of the first subject and the head and limb points of the second subject. Perform a weighted operation on the first similarity and the second similarity corresponding to the head and limb points to obtain a similarity score corresponding to the head and limb points, which is used as a sub-result for evaluating the head and limb points when the first subject performs the target action.
[0155] It should be noted that for the above formulas related to the head limb points, j is a positive integer not exceeding n, and n refers to the number of head limb points.
[0156] Based on the above method, the similarity C between the hand limb points is calculated hand and the similarity C between the leg limb points leg , which are respectively used as sub-results for evaluating the hand limb points and leg limb points when the first subject performs the target action.
[0157] Based on the similarities corresponding to each type of limb key point, a weighted operation is performed to obtain a similarity score, which is used as the complete evaluation result of the target action. The weights corresponding to each type of limb key point can be any real number.
[0158] Exemplarily, as shown in Formula Six below, C overall is the similarity score for comprehensively evaluating the target action. In Formula Six, C head is the similarity corresponding to the head limb points, C hand is the similarity corresponding to the hand limb points, C leg is the similarity corresponding to the leg limb points, and the weights assigned to them are all 1 / 3.
[0159] Formula Six:
[0160]
[0161] In summary, when evaluating the target action performed by the first subject, after obtaining the corresponding sub-data from the reference key point data and the limb key point data for different types of limb points respectively, calculate the similarities in different axis directions and the similarity in depth of the image coordinates of each type of limb key point, and then perform a weighted operation to obtain the similarity between this type of limb key points. After performing a weighted operation on the similarities corresponding to each type of limb key point respectively, a similarity score when the first subject performs the target action is obtained, which is used as the final evaluation result.
[0162] In some embodiments, in addition to the above method of performing similarity analysis by obtaining the three-dimensional information corresponding to the limb key points, other methods can also be used to evaluate the target action, and this embodiment does not limit this.
[0163] In summary, when evaluating the target action performed by the first subject using the action evaluation method provided in this application, the reference action performed by the second subject is used as the evaluation criterion. By obtaining the three-dimensional information of the limb key points of the first subject and the second subject in the image respectively, the matching degree between the reference action and the target action is determined based on the difference between the three-dimensional information, and the evaluation result of the target action is obtained. The three-dimensional information includes the planar coordinates and depth information of the limb key points, which makes the positioning of the limb key points in the three-dimensional space more accurate. Compared with the method of performing action matching and evaluation only based on planar coordinates, it can reduce the influence of action deformation caused by perspective problems and inconsistent longitudinal distances in the planar image, detect the movement trajectory of the limb key points more accurately, and improve the accuracy and comprehensiveness of the action evaluation result.
[0164] Figure 4 FIG. 4 is a flowchart of an action evaluation method provided by another exemplary embodiment of this application, including the following steps.
[0165] Step 410, obtaining frame sequence data based on the reference video.
[0166] Among them, the reference video is a pre-recorded video used as the evaluation benchmark for actions, and is used to display the picture of the second subject performing the reference action.
[0167] The reference video contains the complete process of the second subject performing the reference action. The reference video contains at least two frames of images, and each frame of image represents a video frame corresponding to different timestamps in the reference video.
[0168] Extract the video frames in the reference video to obtain the frame sequence data. The frame sequence data includes at least two frames of images in the reference video and the timestamps respectively corresponding to the at least two frames of images. The timestamps are used to indicate the progress of each frame of image in the reference video.
[0169] Exemplarily, the extracted frame sequence data is a four-dimensional tensor X ∈ R T×C×H×W , where R is a real number, T is the number of video frames, C is the number of color channels of the video frame (including three channels, representing Red, Green, and Blue respectively, that is, the RGB channels), H is the resolution of the reference video in height, and W is the resolution of the reference video in width.
[0170] The number of video frames T included in the reference video is determined by the product of the frame rate fps of the reference video and the video duration s. The timestamp t corresponding to each frame of the picture = i / T = i / fps*s, where i is the serial number of the video frame.
[0171] For example, if the video duration of the reference video is 10 seconds and the video frame rate fps is 24 (24 frames per second), then there are a total of 10 * 24 = 240 video frames. A frame sequence composed of video frames numbered from 1 to 240 is obtained by sorting according to the timestamps of the video frames.
[0172] The timestamp of the first video frame is 0 seconds, corresponding to the start picture of the reference video; the timestamp of the 24th video frame is 1 second, corresponding to the picture at the 1-second progress of the reference video.
[0173] Step 420, obtaining a key point data set of the second subject based on the frame sequence data.
[0174] Among them, the key point data set contains reference key point data. The key point data set contains the image coordinates and depth information corresponding to the second limb key points when the second subject performs a reference action in each frame of the reference video.
[0175] Optionally, key point analysis is performed on each frame of the frame sequence data respectively to obtain key point data corresponding to at least two frames of images.
[0176] Integrate the key point data corresponding to at least two frames of images based on the timestamps corresponding to at least two frames of images to obtain a key point data set.
[0177] Exemplarily, the frame sequence data is input into a pre-trained AI model to extract the key point data of each frame of the image, and a key point data set P t ∈R N×U×V×D is obtained, where N is the number of second limb key points, U and V represent the position coordinates of the second limb key points in the two-dimensional space (video frame), and D represents the depth of the second limb key points.
[0178] Among them, the pre-trained AI model can identify the limb key points of the subject contained in the image and analyze the two-dimensional coordinates and depth information of the limb key points in the image.
[0179] Exemplarily, the AI model contains an equivalent function Landmark(x) that can extract the three-dimensional information of the limb key points contained in the input image. The key point data set P is obtained through the following formula seven t .
[0180] Formula seven:
[0181] P t = Landmark(X, t)
[0182] Among them, X refers to the frame sequence data, and t refers to the timestamp of each frame image in the frame sequence data. When the equivalence function analyzes the frame sequence data to obtain the key point data, the timestamp data is not involved in the calculation and is only used to mark the correspondence between the key point data of each frame image and the image.
[0183] Obtain the key point data set P of the second subject t Then store it in the database of the terminal local as the benchmark for subsequent evaluation of the actions of other subjects.
[0184] Step 430, display the program interface of the first program.
[0185] Among them, the first program is used to provide an action evaluation function for the first subject, and the action evaluation function is used to evaluate the matching degree between the action of the subject to be evaluated (the first subject in this embodiment) and the action in the reference video.
[0186] The action evaluation function of the first program includes the following application scenarios:
[0187] (1) Real-time collect the action pictures of the subject and evaluate them to generate an evaluation result.
[0188] For example, in a live broadcast scenario, the first subject turns on the live broadcast function of the first program, and the camera component of the terminal collects the pictures of the first subject performing actions and displays them on the terminal screen in real time. During the process of the first subject performing actions, evaluate and generate an evaluation result in real time;
[0189] (2) Receive a video or image uploaded by the user / subject that contains the pictures of the subject performing actions, and generate an evaluation result.
[0190] For example, the first subject pre-records a video of performing actions and sends the video to the first program, and the first program analyzes based on the received video and generates an evaluation result.
[0191] Step 440, in response to receiving an action evaluation start operation, play the reference video at the first moment.
[0192] The reference video is used to display the pictures of the second subject performing reference actions, and the reference video contains at least two frames of images.
[0193] Among them, the reference video is played in the specified area in the program interface, and the program interface also includes an area for displaying the action pictures of the first subject.
[0194] Step 450, obtain the first image.
[0195] Among them, the first image is the image collected by the terminal at the second moment, and the first moment is the moment before the second moment.
[0196] The first image includes a first subject that performs a target action with a reference action as the execution benchmark.
[0197] In this embodiment, a scenario where the first subject broadcasts live through a first program and enables an action evaluation function is taken as an example for illustration. At this time, the first program will collect live video frames in real time based on a certain video frame rate, and these live video frames include the first image.
[0198] Step 460: Obtain the limb key point data of the first subject.
[0199] Among them, the limb key point data includes the first image coordinates and the first depth information corresponding to the first limb key points of the first subject.
[0200] Optionally, the process of obtaining the limb key point data of the first image is the same as the process of performing key point analysis on each frame of image in step 420.
[0201] Input the first image into a pre-trained AI model, and the equivalent function Landmark(x) of the AI model extracts the three-dimensional information of the first limb key points included in the first image to obtain the limb key point data Q. Exemplarily, as shown in Formula VIII below.
[0202] Formula VIII:
[0203] Q = Landmark(Y, t cur )
[0204] t cur is the second moment when the first image is collected, and Y is the first image.
[0205] Exemplarily, the limb key point data of the first subject includes the image coordinates and depth information corresponding to 33 first limb key points respectively.
[0206] Among them, it includes 11 head limb points, 12 hand limb points, and 10 leg limb points.
[0207] Step 470: Obtain reference key point data for characterizing the reference action corresponding to the target action.
[0208] Among them, the reference key point data includes the second image coordinates and the second depth information corresponding to the second limb key points of the second subject, and the second subject is the execution subject of the reference action.
[0209] Since the terminal device for collecting the first image and the terminal device for collecting the reference video are not necessarily the same, the video frame rate of the obtained reference video may be inconsistent with the frame rate when the terminal collects the first image in real time. Therefore, it is necessary to determine the video frame in the reference video whose timestamp matches the first image. The key point data corresponding to this video frame is used as the reference key point data.
[0210] Optionally, obtain the time difference between the second moment and the first moment, where the time difference is used to indicate the playback progress of the reference video at the second moment.
[0211] Since the starting playback moment of the reference video is the first moment, and this action evaluation scenario is a live / real-time scenario, it indicates that the action benchmark of the first subject at any moment is the picture played by the reference video at that moment. Therefore, the time difference can determine the time stamp corresponding to the picture played by the reference video at the second moment.
[0212] For example, if the first moment is 10:00:00 and the second moment is 10:00:01, it means that when the first image is captured, the playback progress of the reference video is 1 second.
[0213] Obtain reference key point data from the key point data set based on the time difference and the time stamps included in the key point data set, where the reference key point data corresponds to the second image in at least two frames of images, and the progress of the second image in the reference video matches the playback progress of the reference video at the second moment.
[0214] Exemplarily, refer to Formula Nine below, which is used to determine the value of the sequence number i of the video frame represented by the second image in the frame sequence.
[0215] Formula Nine:
[0216]
[0217] where t diff refers to the absolute time difference between the second moment t cur and the first moment t start , that is, t diff =t cur -t start , i is used to determine the sequence number of the second image in the frame sequence, and fps is the video frame rate of the reference video.
[0218] Formula Nine represents the value of i when the difference between the progress of the second image in the reference video and the playback progress of the reference video at the second moment is minimized, which can solve the error caused by inconsistent frame rates.
[0219] Step 480, determine the evaluation result of the target action based on the limb key point data and the reference key point data.
[0220] Among them, the evaluation result is used to express the matching degree between the target action and the reference action.
[0221] In some embodiments, when obtaining the evaluation result by other means, the evaluation result includes text evaluation content, and the matching degree between the target action and the reference action is described in words.
[0222] Optionally, convert the evaluation result into an evaluation score, which is used to describe the evaluation result in numerical form.
[0223] Display the evaluation score in the result display area of the first program in a first prominent form, and there is a corresponding relationship between the type of the first prominent form and the value of the evaluation score.
[0224] Exemplarily, the evaluation result of the target action includes the following contents: (1) the evaluation sub-result of the head limb points; (2) the evaluation sub-result of the hand limb points; (3) the evaluation sub-result of the leg limb points; (4) the overall evaluation result of the target action.
[0225] After converting the evaluation result into an evaluation score, obtain the evaluation scores corresponding to the above four contents respectively.
[0226] Exemplarily, the result display area of the first program includes display elements corresponding to the above four contents respectively, and the display form of each display element is determined by the evaluation score of its corresponding content.
[0227] Schematically, as Figure 5 shown, the first prominent form means representing the evaluation scores of the four contents with squares of a preset size in the result display area, including a first square 501 representing the evaluation sub-result of the head limb points, a second square 502 representing the evaluation sub-result of the hand limb points, and a third square 503 representing the evaluation sub-result of the leg limb points.
[0228] Among them, there is the following corresponding relationship between the colors of the first square 501, the second square 502, and the third square 503 and the evaluation scores.
[0229] If the evaluation score C of this limb key point k < 0.5, then render this square in the first color;
[0230] If the evaluation score 0.5 < C of this limb key point k < 0.75, then render this square in the second color;
[0231] If the evaluation score C of this limb key point k > 0.75, then render this square in the third color.
[0232] Among them, k ∈ {head head limb points, hand hand limb points, leg leg limb points}.
[0233] For the overall evaluation result, at least one square is displayed in the first area 504. The number of squares is proportional to the evaluation score, and different colors are presented according to the interval to which the evaluation score belongs. Refer to the correspondence between the above evaluation score and the first color, the second color and the third color.
[0234] For example, Figure 5 The diagram includes three images collected at three different times: a first example image 510 , a second example image 520 , and a third example image 530 .
[0235] In the first example image 510, for the actions of the first subject and the second subject, only the evaluation score between the hand limb points is lower than 0.5, and the second block 502 presents the first color; the evaluation scores between the leg limb points and the head limb points are both 0.9, the first block 501 and the third block 503 respectively present the third color, and the overall evaluation score is 0.7, and 6 blocks of the second color are displayed in the first area 504.
[0236] In the second example image 520, for the actions of the first subject and the second subject, only the evaluation score between the leg limb points is lower than 0.5, and the third block 503 presents the first color; the evaluation scores between the hand limb points and the head limb points are both 0.9, the first block 501 and the second block 502 present the third color respectively, and the overall evaluation score is 0.6, and 5 blocks of the second color are displayed in the first area 504.
[0237] In the third example image 530, the evaluation scores between the head limb points, hand limb points and leg limb points in the actions of the first subject and the second subject are all 0.9, the first block 501, the second block 502 and the third block 503 respectively present the third color, the overall evaluation score is 0.9, and 8 blocks presenting the third color are displayed in the first area 504.
[0238] In summary, the action evaluation method provided by this application can solve the out-of-sync problem caused by different frame rates between the reference video and the actually captured image, enable videos shot at different frame rates to display corresponding pictures in the same time dimension, and improve the accuracy of the results obtained when evaluating actions based on the three-dimensional information between video frames. When evaluating the target action performed by the first subject, the reference action performed by the second subject is used as the evaluation criterion. By obtaining the three-dimensional information of the limb key points of the first subject and the second subject in the image respectively, the matching degree between the reference action and the target action is determined based on the difference between the three-dimensional information, and the evaluation result of the target action is obtained. The three-dimensional information includes the planar coordinates and depth information of the limb key points, making the positioning of the limb key points in the three-dimensional space more accurate. Compared with the method of performing action matching and evaluation only based on planar coordinates, it can reduce the influence of action deformation caused by perspective problems and inconsistent longitudinal distances in the planar image, detect the movement trajectory of the limb key points more accurately, and improve the accuracy and comprehensiveness of the action evaluation result.
[0239] In some embodiments, the first program can provide an action follow-along function, display the reference video and the action picture of the first subject on the program interface at the same time, and directly map the action of the second subject in the reference video to the position that coincides with the first subject in the picture, facilitating the first subject to intuitively see the differences in posture, angle, amplitude, etc. between their own execution of the target action and the reference action, and improving the follow-along effect and user experience.
[0240] Figure 6 It is a flowchart of an action follow-along method provided by an exemplary embodiment of this application, including the following steps.
[0241] Step 610, display the program interface of the first program.
[0242] The first program is used to provide an action evaluation function for the first subject, and the action evaluation function is used to evaluate the matching degree between the action of the first subject and the action in the reference video.
[0243] The reference video is a pre-recorded video used as the evaluation benchmark for actions, and is used to display the picture of the second subject performing the reference action.
[0244] Among them, the first program can also provide a live broadcast function. When the first subject turns on the live broadcast function, the picture containing the first subject is collected through the camera component of the terminal, and the picture is displayed in real time through the program interface.
[0245] Step 620, display the live broadcast picture.
[0246] During the live broadcast, the first subject can observe and adjust its own actions through the picture presented on the terminal screen. The first image is used to refer to a frame of the live broadcast picture at the current moment, and the first image contains the first subject.
[0247] Step 630, in response to receiving the action evaluation start operation, play the reference video at the first moment.
[0248] In the case of playing the reference video, the first program evaluates the picture of the first subject performing actions collected in real time by the camera component.
[0249] The reference video is used to show the picture of the second subject performing the reference action. The reference video contains at least two frames of images, and the reference video is the benchmark for evaluating the actions of the first subject.
[0250] That is, at this time, the reference video and the first image are simultaneously displayed on the terminal screen.
[0251] Step 640, in response to receiving the action mapping operation, map the actions of the second subject in the reference video to the first subject in the first image based on the first limb key points and the second limb key points.
[0252] Among them, the limb key point is the key point marked at the specified position in the limb part of the subject. The first limb key point refers to the limb key point of the first subject, and the second limb key point refers to the limb key point of the second subject.
[0253] Exemplarily, determine the second image with a matching timestamp from the reference video based on the moment corresponding to the first image. That is, the target action performed by the first subject in the first image and the reference action performed by the second subject in the second image are based on the same action, so that the actions of the first subject and the second subject in the screen picture are synchronized.
[0254] Perform key point analysis on the first image to determine the three-dimensional coordinates of the first limb key point in the first image (including the two-dimensional image coordinates and depth information with the first image as the plane); perform key point analysis on the second image to determine the three-dimensional coordinates of the second limb key point in the second image (including the two-dimensional image coordinates and depth information with the second image as the plane).
[0255] Map the second limb key point to the first image, and map the three-dimensional coordinates of the second limb key point to the coordinate system same as that of the first limb key point.
[0256] Align the first limb key points with the second limb key points based on the ICP algorithm (Iterative Closest Point algorithm, an algorithm for matching and aligning two point sets), so that the positions of the second limb key points on the first image are as close as possible to the corresponding second limb key points until the preset alignment requirements are met.
[0257] For example, the preset alignment requirement means that within the first image, the distance between the first limb key point and the second limb key point with the same serial number is less than the preset threshold.
[0258] Based on the aligned limb key points, map the actions of the second subject to the first subject and display them in a blurred manner. At this time, in the live broadcast screen of the first program, it shows that the first subject performs action practice following the blurred actions of the second subject that coincides with itself.
[0259] Exemplarily, use a depth estimation model to perform depth estimation on the first image and the second image respectively to obtain the initial depth maps corresponding to the two images. Among them, the initial depth map is used to provide the depth information of the image, indicating the distance relationship between the subject in the image and the camera used when collecting the image, and showing the depth distribution of the content in the image.
[0260] Use a feature extraction algorithm to extract feature points in the first image and the second image respectively. For example, the ORB algorithm (Oriented fast and Rotated BRIEF); and use a matching algorithm to obtain the corresponding relationship between the feature points of the two images. For example, the BFmatcher algorithm (Brute-Force matcher). Obtain a transformation matrix based on the corresponding relationship. The transformation matrix is used to transform the feature points in the second image to the first image so that the feature points are roughly aligned.
[0261] Among them, the feature points can be limb key points or other key points different from the limb key points. This embodiment does not limit this.
[0262] Adjust the initial depth map based on the corresponding relationship between the aligned feature points so that the depths of the corresponding feature points are as consistent as possible to obtain the updated depth map. For example, calculate the depth mean between the feature points representing the same position in the two images and adjust the depth of the feature points in the second image based on the mean.
[0263] After completing the above depth matching, transform the two-dimensional coordinates of the second limb key points in the second image to the two-dimensional coordinate system with the first image as the plane to achieve the matching of the two-dimensional coordinates.
[0264] Based on the above example, it is possible to map the three-dimensional coordinates of the key points of the second limb to the same coordinate system as the three-dimensional coordinates of the key points of the first limb, completing the action mapping from the second subject to the first subject.
[0265] In summary, the action following method provided in this application can, through the live broadcast screen, compare the limb movements of the first subject with the reference movements executed by the second subject in the reference video in real time, accurately correct errors, improve the efficiency of the user's action following, and enhance the interactive experience of the first subject during action following.
[0266] Figure 7 It is a structural block diagram of an action evaluation device provided by an exemplary embodiment of this application, as Figure 7 shown, the device includes the following parts.
[0267] An image acquisition module 710, configured to acquire a first image, where the first image includes a first subject performing a target action;
[0268] A data acquisition module 720, configured to determine limb key point data for characterizing the target action based on the first image, where the limb key point data includes a first image coordinate and a first depth information corresponding to a first limb key point of the first subject;
[0269] The data acquisition module 720 is further configured to acquire reference key point data for characterizing a reference action corresponding to the target action, where the reference key point data includes a second image coordinate and a second depth information corresponding to a second limb key point of a second subject, and the second subject is the executor of the reference action;
[0270] An evaluation module 730, configured to determine an evaluation result of the target action based on the limb key point data and the reference key point data.
[0271] In an optional embodiment, the evaluation module 730 is further configured to perform a similarity analysis on the limb key point data and the reference key point data, and obtain a similarity score as the evaluation result of the target action.
[0272] In an optional embodiment, the evaluation module 730 is further configured to obtain a first similarity between the first image coordinate and the second image coordinate; obtain a second similarity between the first depth information and the second depth information; and obtain the similarity score based on a weighted sum between the first similarity and the second similarity.
[0273] In an alternative embodiment, the evaluation module 730 is further configured to obtain a first coordinate similarity between a first coordinate of the first image coordinate in a first axis direction and a first coordinate of the second image coordinate in the first axis direction; and obtain a second coordinate similarity between a second coordinate of the first image coordinate in a second axis direction and a second coordinate of the second image coordinate in the second axis direction; and obtain the first similarity based on a weighted sum between the first coordinate similarity and the second coordinate similarity.
[0274] In an alternative embodiment, the limb key points include at least two types of key points representing different limb parts; the i-th type of key point in the limb key points corresponds to the i-th limb part of the first subject and the second subject, where i is a positive integer;
[0275] The evaluation module 730 is further configured to, for the i-th type of key point, obtain an evaluation sub-result for the i-th limb part based on a difference between the i-th sub-data in the limb key point data and the i-th reference sub-data in the reference key point data; and obtain the evaluation result for the target action based on the evaluation sub-results respectively corresponding to the at least two types of key points.
[0276] In an alternative embodiment, as Figure 8 shown, the apparatus further includes:
[0277] A display module 740, configured to display a program interface of a first program, where the first program is used to provide an action evaluation function for the first subject; in response to receiving an action evaluation start operation, play a reference video at a first moment, where the reference video is used to show a picture of the second subject performing the reference action, and the reference video includes at least two frames of images; where the first image is an image collected by the terminal at a second moment, and the first moment is a moment before the second moment.
[0278] In an alternative embodiment, the data acquisition module 720 is further configured to obtain frame sequence data based on the reference video, where the frame sequence data includes the at least two frames of images in the reference video and time stamps respectively corresponding to the at least two frames of images, and the time stamps are used to indicate the progress of each frame of image in the reference video; and obtain a key point data set of the second subject based on the frame sequence data, where the key point data set includes the reference key point data, and the key point data set includes image coordinates and depth information corresponding to the second limb key points when the second subject performs the reference action in each frame of image in the reference video.
[0279] In an alternative embodiment, the data acquisition module 720 is further configured to perform key point analysis on each frame of image in the frame sequence data to obtain key point data corresponding to each of the at least two frames of images; and integrate the key point data corresponding to each of the at least two frames of images based on the timestamps corresponding to each of the at least two frames of images to obtain the key point data set.
[0280] In an alternative embodiment, the data acquisition module 720 is further configured to obtain a time difference between the second moment and the first moment, where the time difference is used to indicate the playback progress of the reference video at the second moment; and obtain the reference key point data from the key point data set based on the time difference and the timestamps included in the key point data set, where the reference key point data corresponds to the second image among the at least two frames of images, and the progress of the second image in the reference video matches the playback progress of the reference video at the second moment.
[0281] In an alternative embodiment, the evaluation module 730 is further configured to convert the evaluation result into an evaluation score, where the evaluation score is used to describe the evaluation result in numerical form;
[0282] The display module 740 is further configured to display the evaluation score in a first prominent form in the result display area of the first program, where there is a corresponding relationship between the type of the first prominent form and the value of the evaluation score.
[0283] In an alternative embodiment, the device further includes:
[0284] A mapping module 750, configured to, in response to receiving an action mapping operation, map the action of the second subject in the reference video to the first subject in the first image based on the first limb key points and the second limb key points, where the first limb key points and the second limb key points meet the preset mapping coincidence requirements.
[0285] In summary, when the action evaluation device provided in this application evaluates the target action performed by the first subject, it uses the reference action performed by the second subject as the evaluation criterion. By obtaining the three-dimensional information of the limb key points of the first subject and the second subject in the image respectively, it determines the matching degree between the reference action and the target action based on the difference between the three-dimensional information, and obtains the evaluation result of the target action. The three-dimensional information includes the planar coordinates and depth information of the limb key points, which makes the positioning of the limb key points in the three-dimensional space more accurate. Compared with the method of performing action matching and evaluation only based on planar coordinates, it can reduce the influence of action deformation caused by perspective problems and inconsistent longitudinal distances in the planar image, more accurately detect the movement trajectory of the limb key points, and improve the accuracy and comprehensiveness of the action evaluation result.
[0286] It should be noted that: for the action evaluation device provided in the above embodiment, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the action evaluation device provided in the above embodiment and the action evaluation method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0287] Figure 9 The block diagram of a computer device 900 provided by an exemplary embodiment of this application is shown. The computer device 900 can be: a smart phone, a tablet computer, a Moving Picture Experts Group Audio Layer III player (MP3), a Moving Picture Experts Group Audio Layer IV (MP4) player, a notebook computer or a desktop computer. The computer device 900 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0288] Generally, the computer device 900 includes: a processor 901 and a memory 902.
[0289] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 901 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0290] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 901 to implement the evaluation method of the actions provided in the method embodiments of the present application.
[0291] In some embodiments, the computer device 900 further includes some other components 903, and the type and quantity of the other components 903 may be selected based on the functional requirements of the computer device 900. Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the computer device 900, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0292] Optionally, the computer-readable storage medium may include: Read Only Memory (ROM), Random Access Memory (RAM), Solid State Drives (SSD), or optical discs, etc. Among them, the random access memory may include Resistance Random Access Memory (ReRAM) and Dynamic Random Access Memory (DRAM). The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0293] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the evaluation method of any one of the actions described in the embodiments of the present application above.
[0294] An embodiment of the present application further provides a computer-readable storage medium. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the evaluation method of any one of the actions described in the embodiments of the present application above.
[0295] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the evaluation method of any one of the actions described in the above embodiments.
[0296] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0297] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for evaluating an action, characterized in that, The method includes: Obtaining a first image, where the first image includes a first subject performing a target action; Determining limb key-point data for characterizing the target action based on the first image, where the limb key-point data includes a first image coordinate and a first depth information corresponding to a first limb key-point of the first subject; Obtaining reference key-point data for characterizing a reference action corresponding to the target action, where the reference key-point data includes a second image coordinate and a second depth information corresponding to a second limb key-point of a second subject, and the second subject is the performer of the reference action; Determining an evaluation result of the target action based on the limb key-point data and the reference key-point data.
2. The method according to claim 1, wherein The determining the evaluation result of the target action based on the limb key-point data and the reference key-point data includes: Performing a similarity analysis on the limb key-point data and the reference key-point data to obtain a similarity score as the evaluation result of the target action.
3. The method according to claim 2, wherein The performing a similarity analysis on the limb key-point data and the reference key-point data to obtain a similarity score as the evaluation result of the target action includes: Obtaining a first similarity between the first image coordinate and the second image coordinate; Obtaining a second similarity between the first depth information and the second depth information; Obtaining the similarity score based on a weighted sum between the first similarity and the second similarity.
4. The method according to claim 3, wherein The obtaining the first similarity between the first image coordinate and the second image coordinate includes: Obtaining a first coordinate similarity between a first coordinate of the first image coordinate in a first axis direction and a first coordinate of the second image coordinate in the first axis direction; and obtaining a second coordinate similarity between a second coordinate of the first image coordinate in a second axis direction and a second coordinate of the second image coordinate in the second axis direction; Obtaining the first similarity based on a weighted sum between the first coordinate similarity and the second coordinate similarity.
5. The method according to any one of claims 1 to 4, characterized in that The limb key-points include at least two types of key-points representing different limb parts; The i-th type of key-point in the limb key-points corresponds to the i-th limb part of the first subject and the second subject, where i is a positive integer; The determining the evaluation result of the target action based on the limb key-point data and the reference key-point data includes: For the i-th type of key-point, obtaining an evaluation sub-result of the i-th limb part based on a difference between an i-th sub-data in the limb key-point data and an i-th reference sub-data in the reference key-point data; Obtaining the evaluation result of the target action based on the evaluation sub-results corresponding to the at least two types of key-points respectively.
6. The method according to any one of claims 1 to 4, characterized in that Before obtaining the first image, it further includes: Displaying a program interface of a first program, where the first program is used to provide an action evaluation function for the first subject; In response to receiving an action evaluation start operation, playing a reference video at a first moment, where the reference video is used to show a picture of the second subject performing the reference action, and the reference video includes at least two frames of images; Wherein, the first image is an image acquired by the terminal at a second moment, and the first moment is a moment before the second moment.
7. The method according to claim 6, characterized in that, Before obtaining the reference key point data for characterizing the reference action corresponding to the target action, it further includes: Obtaining frame sequence data based on the reference video, where the frame sequence data includes the at least two images in the reference video and the timestamps respectively corresponding to the at least two images, and the timestamps are used to indicate the progress of each frame image in the reference video; Obtaining a set of key point data of the second subject based on the frame sequence data, where the set of key point data includes the reference key point data, and the set of key point data includes the image coordinates and depth information corresponding to the second limb key points when the second subject performs the reference action in each frame image of the reference video.
8. The method according to claim 7, characterized in that, The obtaining the set of key point data of the second subject based on the frame sequence data includes: Performing key point analysis on each frame image in the frame sequence data respectively to obtain key point data corresponding to the at least two images respectively; Integrating the key point data corresponding to the at least two images respectively based on the timestamps corresponding to the at least two images respectively to obtain the set of key point data.
9. The method according to claim 7, wherein The obtaining the reference key point data for characterizing the reference action corresponding to the target action includes: Obtaining the time difference between the second moment and the first moment, where the time difference is used to indicate the playing progress of the reference video at the second moment; Obtaining the reference key point data from the set of key point data based on the time difference and the timestamps included in the set of key point data, where the reference key point data corresponds to the second image among the at least two images, and the progress of the second image in the reference video matches the playing progress of the reference video at the second moment.
10. The method according to claim 6, wherein After determining the evaluation result of the target action based on the limb key point data and the reference key point data, it further includes: Converting the evaluation result into an evaluation score, where the evaluation score is used to describe the evaluation result in numerical form; Displaying the evaluation score in the result display area of the first program in a first prominent form, and there is a corresponding relationship between the type of the first prominent form and the value of the evaluation score.
11. The method according to claim 6, wherein The method further includes: In response to receiving an action mapping operation, mapping the action of the second subject in the reference video to the first subject in the first image based on the first limb key points and the second limb key points, where the first limb key points and the second limb key points meet the preset mapping coincidence requirements.
12. An evaluation device for an action, characterized in that, The device includes: An image acquisition module, configured to acquire a first image, where the first image includes a first subject performing a target action; A data acquisition module, configured to determine limb key point data for characterizing the target action based on the first image, where the limb key point data includes the first image coordinates and the first depth information corresponding to the first limb key points of the first subject. The data acquisition module is further configured to acquire reference key-point data for characterizing a reference action corresponding to the target action, where the reference key-point data includes second image coordinates and second depth information corresponding to second limb key-points of a second subject, and the second subject is the execution subject of the reference action; An evaluation module, configured to determine an evaluation result of the target action based on the limb key-point data and the reference key-point data.
13. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the action evaluation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the action evaluation method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the action evaluation method according to any one of claims 1 to 11.
Citation Information
Cited By
Video generation method, system and model
CN122073636A