Method and device for determining category of marker points for motion capture, and terminal

By acquiring multiple images and inputting the trained classification model, the existing motion capture methods are solved in terms of accuracy and efficiency, and the accurate and efficient determination of the movement posture of real actors is achieved.

CN113989923BActive Publication Date: 2025-06-03MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111212197.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-06-03
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

The existing motion capture methods have shortcomings in accuracy and efficiency, and it is difficult to accurately and efficiently determine the movement posture of real actors.

Method used

By acquiring multiple images, the position information of the marking points to be identified is determined, and this information is input into the trained classification model to obtain the category information of the marking points. The classification model uses the training marking point set to train the preset model, including the position information and real category information of multiple sample marking points under multiple action postures.

Benefits of technology

This method can accurately and efficiently determine the subject's motion posture, avoid the problem of error accumulation, improve the accuracy of marking point categories, and improve the efficiency of motion capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989923B_ABST
    Figure CN113989923B_ABST
Patent Text Reader

Abstract

A method and device and terminal for determining the category of marker points for motion capture. The method includes: obtaining multiple images at the current moment, where each image includes multiple marker points to be recognized; determining the position information of the multiple marker points to be recognized at the current moment based on the multiple images; inputting the position information of the multiple marker points to be recognized into a classification model to obtain the category information of the multiple marker points to be recognized, where the classification model is obtained by training a preset model using a training marker point set, and the training marker point set includes the position information and true category information of multiple sample marker points under multiple action postures. Through the solution of the present invention, it is beneficial to accurately and efficiently determine the action posture of the object being photographed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for determining the category of marker points for motion capture, and a terminal. Background Art

[0002] Motion capture technology is a technology that uses sensors to obtain the motion data of real actors for recording and migrates it to virtual characters. Currently, it plays a crucial role in fields such as movie special effects, game production, and virtual live broadcast. However, the accuracy and efficiency of existing motion capture methods need to be improved. Therefore, there is an urgent need for an improved solution that can help accurately and efficiently determine the motion postures of real actors. Summary of the Invention

[0003] The technical problem solved by the present invention is to provide a method and device for determining the category of marker points for motion capture, and a terminal, so as to accurately and efficiently determine the motion postures of the object to be photographed.

[0004] To solve the above technical problem, an embodiment of the present invention provides a method for determining the category of marker points for motion capture, the method includes: obtaining multiple images at the current moment, where each image includes multiple marker points to be recognized; determining the position information of the multiple marker points to be recognized at the current moment according to the multiple images; inputting the position information of the multiple marker points to be recognized into a classification model to obtain the category information of the multiple marker points to be recognized, where the classification model is obtained by training a preset model with a training marker point set, and the training marker point set includes the position information and real category information of multiple sample marker points in multiple motion postures.

[0005] Optionally, the method further includes: determining the motion posture of the object to be photographed at the current moment according to the category information of the multiple marker points to be recognized.

[0006] Optionally, the position information of the multiple marker points to be recognized is a coordinate matrix corresponding to the multiple marker points to be recognized at the current moment. Determining the position information of the multiple marker points to be recognized according to the multiple images includes: determining an initial coordinate matrix corresponding to the multiple marker points to be recognized according to the position of each marker point to be recognized in each image; processing the initial coordinate matrix according to the centroid coordinates and / or rotation matrix to obtain a processed coordinate matrix; where the centroid coordinates are the position of the centroid of the object to be photographed at the current moment in the space coordinate system, and the rotation matrix is obtained by processing the initial coordinate matrix using the principal component analysis algorithm.

[0007] Optionally, the classification model includes a first fully-connected layer, a feature fusion module, and a first normalization module. Among them, the first fully-connected layer is used to calculate the feature information of each to-be-recognized marker point according to the position information of the to-be-recognized marker point; the feature fusion module is used to update the feature information of the to-be-recognized marker point according to the feature information of multiple to-be-recognized marker points within the neighborhood of the to-be-recognized marker point; the first normalization module is used to determine the category information of the multiple to-be-recognized marker points according to the feature information of the multiple to-be-recognized marker points.

[0008] Optionally, the feature fusion module includes a geometric information fusion sub-module and / or an edge convolution sub-module: The geometric information fusion sub-module is used to update the feature information of each to-be-recognized marker point according to the feature information of the first adjacent marker points of the to-be-recognized marker point, where the first adjacent marker points are multiple to-be-recognized marker points within the geometric neighborhood of the to-be-recognized marker point; the edge convolution sub-module is used to update the feature information of each to-be-recognized marker point according to the feature information of the second adjacent marker points of the to-be-recognized marker point, where the second adjacent marker points are multiple to-be-recognized marker points within the feature neighborhood of the to-be-recognized marker point.

[0009] Optionally, the position information of the multiple to-be-recognized marker points is a multi-dimensional coordinate matrix, the category information is a probability matrix, and the classification model includes: a stretching processing module, a second fully-connected layer, and a second normalization module. Among them, the stretching processing module is used to perform stretching processing on the multi-dimensional coordinate matrix to obtain a stretched coordinate matrix, where the stretched coordinate matrix is a one-dimensional matrix; the second fully-connected layer is used to process the stretched coordinate matrix to obtain a one-dimensional output matrix; the second normalization module is used to process the output matrix to obtain the probability matrix.

[0010] Optionally, the category information of the multiple to-be-recognized marker points is an initial probability matrix, and the initial probability matrix includes the probabilities of each to-be-recognized marker point belonging to each preset category. The preset categories include the ghost point category. The method further includes: determining a predicted ghost point marker, where the predicted ghost point marker is the to-be-recognized marker point with the highest probability of belonging to the ghost point category among the probabilities of belonging to each preset category; removing the information of the predicted ghost point marker and the information of the ghost point category from the initial probability matrix to obtain a first probability matrix; determining the preset category corresponding to each to-be-recognized marker point according to the first probability matrix.

[0011] Optionally, determining the preset category corresponding to each to-be-recognized marked point according to the first probability matrix includes: processing the first probability matrix to obtain a second probability matrix, where the second probability matrix is a doubly stochastic matrix; determining the preset category corresponding to each to-be-recognized marked point according to the second probability matrix.

[0012] Optionally, determining the preset category corresponding to each to-be-recognized marked point according to the second probability matrix includes: for each to-be-recognized marked point in the second probability matrix, determining the maximum value among the probabilities that the to-be-recognized marked point belongs to each preset category in the second probability matrix, and taking the preset category corresponding to the maximum value as the preset category corresponding to the to-be-recognized marked point.

[0013] Optionally, determining the preset category corresponding to each to-be-recognized marked point according to the first probability matrix includes: determining the cost matrix of the first probability matrix; determining the preset category corresponding to each to-be-recognized marked point in the first probability matrix according to the bipartite matching algorithm and the cost matrix.

[0014] Optionally, the method for training the classification model includes: obtaining the training marked point set; inputting the position information of the multiple sample marked points into the preset model to obtain predicted category information; calculating a loss value according to the predicted category information and the true category information; updating the preset model according to the loss value, and when a preset stop condition is satisfied, obtaining the classification model.

[0015] Optionally, the multiple sample marked points include ghost points and multiple sample entity marked points, the predicted category information is an initial predicted probability matrix, the initial predicted probability matrix includes the probabilities that each sample marked point belongs to each preset category, the preset category includes the ghost point category, and calculating the loss value according to the predicted category information and the true category information includes: removing from the initial predicted probability matrix the probabilities that the ghost points belong to each preset category and the probabilities that each sample marked point belongs to the ghost point category to obtain a first predicted probability matrix; calculating the loss value according to the first predicted probability matrix and the true category information.

[0016] Optionally, calculating the loss value according to the first predicted probability matrix and the true category information includes: processing the first predicted probability matrix to obtain a second predicted probability matrix, where the second predicted probability matrix is a doubly stochastic matrix; calculating the loss value according to the second predicted probability matrix and the true category information.

[0017] Optionally, the position information of the multiple sample marker points is a sample coordinate matrix. Calculating the loss value according to the first prediction probability matrix and the true class information includes: removing the coordinates of the ghost points from the sample coordinate matrix to obtain a sample coordinate matrix after removal; calculating a prediction distance matrix according to the sample coordinate matrix after removal and the first prediction probability matrix, where the prediction distance matrix includes the predicted distances between every two sample entity marker points; calculating the loss value according to the prediction distance matrix and the true distance matrix.

[0018] Optionally, before inputting the position information of the multiple sample marker points into the preset model, the method further includes: performing enhancement processing on the training marker point set, where the enhancement processing includes one or more of the following: performing global rotation and / or global translation processing on the position information of at least one group of sample marker points; adding information of ghost points; removing information of at least one sample marker point from the training marker point set; adjusting the position information of at least one sample marker point in the training marker point set; adding information of multiple sample marker points of the subject in a new pose; adding information of multiple sample marker points of other objects other than the subject in the multiple action poses.

[0019] To solve the above technical problems, an embodiment of the present invention further provides a device for determining the class of marker points for motion capture. The device includes: an image acquisition module, configured to acquire multiple images at the current moment, where each image includes multiple marker points to be recognized; a position acquisition module, configured to determine the position information of the multiple marker points to be recognized at the current moment according to the multiple images; a class prediction module, configured to input the position information of the multiple marker points to be recognized into a classification model to obtain the class information of the multiple marker points to be recognized; where the classification model is obtained by training a preset model with training data, and the training data includes the position information and true class information of multiple sample marker points in multiple action poses.

[0020] An embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method are executed.

[0021] An embodiment of the present invention further provides a terminal, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, the steps of the above method are executed.

[0022] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:

[0023] In the solution of the embodiment of the present invention, multiple images at the current moment are obtained, and the position information of multiple to-be-recognized marker points at the current moment is determined according to the multiple images at the current moment. Then, the position information of the multiple to-be-recognized marker points is input into a classification model to obtain the category information of each to-be-recognized marker point. When adopting such a solution, since the classification model is obtained by training a preset model with a training marker point set, and the training marker point set includes the position information and true category information of multiple sample marker points in multiple action postures, therefore, the classification model can be used to determine the category information of each marker point according to the position information of each marker point.

[0024] Compared with the method of determining the category of marker points for motion capture based on a tracking algorithm in the prior art, the solution in the embodiment of the present invention can determine the category information of each to-be-recognized marker point at the current moment without relying on the category information of the marker points in the previous frame image. Thus, the problem of error accumulation can be avoided, the accuracy of the category of the marker points is improved, and thus it is beneficial to accurately determine the action posture.

[0025] In addition, if a shooting interruption occurs, since the solution in the embodiment of the present invention does not rely on the category information of the marker points in the previous frame image and does not require the object to be photographed to return to the initial posture for reshooting, it is beneficial to improve the efficiency of determining the category of the marker points, thereby improving the efficiency of motion capture.

[0026] Furthermore, in the solution of the embodiment of the present invention, preprocessing the initial coordinate matrix according to the centroid coordinates can reduce the influence of the overall position of the object to be photographed on the category information, which is beneficial to improving the accuracy of the category of the subsequent to-be-recognized marker points.

[0027] Furthermore, in the solution of the embodiment of the present invention, normalizing the initial coordinate matrix according to the rotation matrix can reduce the influence of the orientation of the object to be photographed on the category information, which is beneficial to improving the accuracy of the category of the subsequent to-be-recognized marker points.

[0028] Furthermore, in the solution of the embodiment of the present invention, by setting a geometric information fusion sub-module, the feature information of adjacent points of the to-be-recognized marker points in the Euclidean space can be fused, and the local feature information near each to-be-recognized marker point can be stably obtained. Thus, the perception domain of the classification model can be expanded, which is beneficial to improving the accuracy of determining each to-be-recognized marker point.

[0029] Furthermore, in the solution of the embodiment of the present invention, by setting an edge convolution sub-module, the feature information of adjacent points of the to-be-recognized marker points in the feature space can be fused. When the feature information of multiple to-be-recognized marker points changes, the feature neighborhood of each to-be-recognized marker point also changes. Thus, the perception domain of the classification model can be greatly expanded, which is beneficial to improving the accuracy of determining the category of each to-be-recognized marker point.

[0030] Further, in the solution of the embodiment of the present invention, the classification model may include: a stretching processing module, a second fully connected layer, and a second normalization module. Such a classification model is relatively lightweight and is conducive to improving the efficiency of determining the category of the to-be-identified marker points.

[0031] Further, in the solution of the embodiment of the present invention, the probability matrix can be processed to obtain a second probability matrix, and then the matching relationship between each to-be-identified marker point and the preset category can be determined according to the second probability matrix. Since the second probability matrix is a doubly stochastic matrix, the values in the second probability matrix can also be used to indicate the probability that each preset category belongs to each to-be-identified marker point. Therefore, determining the matching relationship between each to-be-identified marker point and the preset category according to the second probability matrix can avoid the situation that multiple to-be-identified marker points belong to the same preset category, which is conducive to improving the accuracy of the category information of the to-be-identified marker points.

[0032] Further, in the solution of the embodiment of the present invention, the cost matrix of the probability matrix is determined, and then the preset category corresponding to each to-be-identified marker point is determined according to the bipartite matching algorithm and the cost matrix. Adopting such a solution is beneficial to avoiding the situation that the to-be-identified marker points belonging to ghost points are misjudged as entity marker points, thereby being beneficial to further improving the accuracy of determining the category information of each to-be-identified marker.

[0033] Further, in the solution of the embodiment of the present invention, the loss value is calculated according to the predicted distance matrix and the true distance matrix. Wherein, the predicted distance matrix includes the predicted distances between every two sample marker points, and each element in the true distance matrix is used to indicate the true distance between the marker point corresponding to the row to which the element belongs and the marker point corresponding to the column to which the element belongs. Since the parts of the object to be photographed usually have the property of a rigid body, the distances between the marker points set on the same part usually do not change with the change of the action posture. Therefore, training the preset model by adopting such a solution can use the rigid body distance between every two marker points as a constraint for supervised learning, which is conducive to improving the accuracy of the model. Description of the Drawings

[0034] Figure 1 is a schematic diagram of an application scenario of a method for determining the category of marker points for motion capture in an embodiment of the present invention;

[0035] Figure 2 is a schematic flowchart of a method for determining the category of marker points for motion capture in an embodiment of the present invention;

[0036] Figure 3 is a schematic structural diagram of a classification model in an embodiment of the present invention;

[0037] Figure 4It is a schematic structural diagram of another classification model in an embodiment of the present invention;

[0038] Figure 5 It is a schematic flowchart of a method for training a classification model in an embodiment of the present invention;

[0039] Figure 6 It is a schematic structural diagram of a device for determining the category of a marker point for motion capture in an embodiment of the present invention;

[0040] Figure 7 It is a front view schematic diagram of a photographed object and a marked object in an embodiment of the present invention;

[0041] Figure 8 It is a rear view schematic diagram of a photographed object and a marked object in an embodiment of the present invention. Detailed implementation manners

[0042] As described in the background art, there is an urgent need for an improved solution that can help accurately and efficiently determine the motion posture.

[0043] The inventors of the present invention have found through research that the existing methods for determining the category of marker points for motion capture are usually optical methods for determining the category of marker points for motion capture. Specifically, a plurality of marker points are pre-placed on a photographed object (such as a real actor, etc.) and an image of the photographed object is collected, and then the part corresponding to the marker point is determined by determining the category of each marker point in the image, whereby the motion posture of the photographed object can be determined. Therefore, determining the category of each marker point in the image is a key step in motion capture.

[0044] In the prior art, a tracking algorithm is usually used to determine the category of each marker point. Specifically, the photographed object makes a specific posture (such as a T-shaped posture or an A-shaped posture) in advance and the photographed object is photographed to obtain an image at the initial moment. The marker points in the image at the initial moment obtained by photographing are compared with the marker points with category labels on a standard template, so as to determine the category of each marker point at the initial moment. The marker points in the image at a subsequent moment will inherit the category of each marker point in the image at the initial moment. Specifically, for each marker point in the image at the current moment, according to the position of the marker point and the positions of each marker point in the image at the previous moment, the marker point in the image at the previous moment that matches the marker point is determined, and the category of the marker point that matches in the image at the previous moment is used as the category of the marker point.

[0045] There are the following problems with such a solution:

[0046] (1) It is easy to cause errors in determining the category of marker points due to missing points or ghost points. Specifically, during the shooting process, due to environmental occlusion and other reasons, some marker points may not be captured, and such points are called "missing points". In addition, due to the possible existence of marker points in the scene that are not attached to the object being photographed, or due to factors such as sensor noise, additional marker points will be recorded, and such points are called "ghost points". The existence of "missing points" and "ghost points" will affect the motion capture process based on the tracking algorithm, resulting in errors in the analysis of the marker point category, causing errors in the motion pose analysis, and making the motion of the virtual object driven by the motion pose appear visibly distorted to the naked eye.

[0047] (2) If there is an error in the category of marker points at any moment, this error will be transmitted to the next frame along with the time-domain information of the tracking algorithm. Specifically, since the category of marker points in the subsequent frame only depends on the category of marker points in the previous frame, the errors in the category of marker points will accumulate over time.

[0048] (3) Once the marker point tracking fails, it cannot be restored immediately. The object being photographed needs to return to a specific pose when the initial frame image was taken for re-shooting or manually correct the category of the marker points, which requires the user to pay a high time and labor cost. Specifically, when the number of "missing points" and "ghost points" is large, or the actions of the object being photographed change violently in a short period of time, resulting in a large distance between the corresponding marker points in adjacent frames, the tracking will fail. After the tracking fails, the object being photographed needs to return to the initial pose to start shooting again; moreover, the results obtained from the shooting usually need to be manually corrected, resulting in unnecessary additional costs.

[0049] To solve the above technical problems, an embodiment of the present invention provides a method for determining the category of marker points for motion capture. To make the above objects, features, and beneficial effects of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings.

[0050] Refer to Figure 1 , Figure 1 is a schematic diagram of an application scenario of a method for determining the category of marker points for motion capture in an embodiment of the present invention. As Figure 1 shown, multiple photographing devices 11 can be set to photograph the object being photographed 12, and the multiple photographing devices 11 can photograph the object being photographed 12 from different perspectives respectively. Among them, the object being photographed 12 is the object photographed by the multiple photographing devices 11. The object being photographed 12 can be a living body that can make autonomous actions, such as a real actor or an animal, etc., or an object that makes corresponding actions according to the received control instructions, such as a robot, etc., but is not limited thereto.

[0051] It should be noted that the number of photographing devices 11 is greater than or equal to 2. More specifically, multiple photographing devices 11 are arranged around the object 12 to be photographed, so as to photograph the object 12 to be photographed in all directions and at all angles as much as possible. In a specific example, multiple photographing devices 11 are infrared cameras, and the images obtained by photographing are binary images, but it is not limited thereto.

[0052] Furthermore, a plurality of marking objects ( Figure 1 not shown) are provided at multiple positions on the object 12 to be photographed. Among them, each part can be a rigid body, and the multiple marking objects at different parts can be arranged in different shapes.

[0053] Referring to Figure 7 and Figure 8 , Figure 7 is a front schematic view of an object 12 to be photographed and a marking object 10 in an embodiment of the present invention, Figure 8 is a back schematic view of an object 12 to be photographed and a marking object 10 in an embodiment of the present invention. More specifically, Figure 7 and Figure 8 show the positions of multiple marking objects 10 on the object 12 to be photographed. The object 12 to be photographed can be a real actor, and multiple marking objects 10 can be pasted on multiple parts of the real actor (for example: head, forearm, upper arm, calf, thigh, upper torso, lower torso, hands and feet, etc.), but it is not limited thereto. It should be noted that the embodiment of the present invention does not limit the setting manner of the marking object 10 on the object 12 to be photographed.

[0054] Continuing to refer to Figure 1 , each image captured by each photographing device 11 may include a plurality of marking points. Among them, the plurality of marking points include a plurality of physical marking points, and the physical marking points are the images of the marking objects provided on the object 12 to be photographed in the image, and the physical marking points are the 2D projections of the corresponding marking objects in the image. That is, each physical marking point in the image has a uniquely corresponding marking object. In an ideal situation, each marking object also has a uniquely corresponding physical marking point, but in the process of photographing, the situation of missing marking objects may occur. At this time, for the missing marking objects, no physical marking points corresponding to the marking objects can be found in the image.

[0055] In addition, the plurality of marking points may further include ghost points, and the ghost points refer to the marking points that do not have corresponding marking objects. In other words, the ghost points refer to the marking points caused by various external factors such as noise, and there is no corresponding marking object on the object 12 to be photographed.

[0056] Further, each photographing device 11 can be connected to the terminal 13. The terminal 13 can obtain images from each photographing device 11 and determine the category information of each marker point in the multiple images captured at the current moment and the action posture of the object 12 to be photographed at the current moment according to the obtained images.

[0057] Referring to Figure 2 , Figure 2 FIG. is a schematic flowchart of a method for determining the category of a marker point for motion capture in an embodiment of the present invention. The method for determining the category of a marker point in the embodiment of the present invention can be used in the field of motion capture. Motion capture can be used in the field of performance animation, that is, the action output of a virtual character can be realized in the same picture. The method can be executed by a terminal, and the terminal can be various terminal devices with data receiving and processing capabilities. For example, it can be a mobile phone, a computer, a tablet computer, etc. The embodiment of the present invention does not limit this. In a specific example, the terminal can be the terminal 13 shown in Figure 1 , but not limited thereto. Through the solution of the embodiment of the present invention, the category information of multiple marker points to be recognized at the current moment and the action posture of the object to be photographed at the current moment can be accurately and efficiently determined only according to the multiple images at the current moment. Figure 2 The method for determining the category of a marker point for motion capture shown in FIG. may include the following steps:

[0058] Step S201: Obtain multiple images at the current moment, where each image includes multiple marker points to be recognized;

[0059] Step S202: Determine the position information of the multiple marker points to be recognized at the current moment according to the multiple images;

[0060] Step S203: Input the position information of the multiple marker points to be recognized into a classification model to obtain the category information of the multiple marker points to be recognized.

[0061] It can be understood that in specific implementation, the method can be implemented in the form of a software program, and the software program runs in a processor integrated inside a chip or a chip module; or, the method can be implemented in a hardware or a combination of hardware and software manner.

[0062] In the specific implementation of step S201, multiple images at the current moment can be obtained, and each image includes multiple marker points to be recognized. Specifically, the multiple images can be images of the object being photographed at the current moment. The multiple images are respectively images of the object being photographed collected by each photographing device at the current moment, and each image includes multiple marker points to be recognized. Among them, the multiple marker points to be recognized can include multiple physical marker points and can also include ghost points. The number of multiple images at the current moment is greater than or equal to 2. For more content about the multiple images and marker points to be recognized at the current moment, reference can be made to the relevant description about Figure 1 and will not be elaborated here.

[0063] In the specific implementation of step S202, an initial coordinate matrix corresponding to multiple marker points to be recognized at the current moment can be determined according to the positions of each marker point to be recognized in each image at the current moment. Among them, the initial coordinate matrix includes the coordinates of each marker point to be recognized in the spatial coordinate system at the current moment.

[0064] Specifically, the multiview 3D reconstruction method can be used to determine the coordinates of each marker point to be recognized in the spatial coordinate system according to the positions of each marker point to be recognized in each image, and thus the initial coordinate matrix can be obtained.

[0065] More specifically, the positions of the marker points to be recognized in each image can be two-dimensional coordinates based on the image coordinate system, and the initial coordinate matrix can include three-dimensional coordinates of multiple marker points in the spatial coordinate system. For example, the initial coordinate matrix can be an N×3 matrix, where N is the number of marker points to be recognized and N is a positive integer. In a specific example, the spatial coordinate system is the world coordinate system, but it is not limited to this.

[0066] Furthermore, the initial coordinate matrix can be preprocessed to improve the accuracy of subsequent determination of the category information of the marker points to be recognized.

[0067] In the first specific example, the centroid coordinates of the object being photographed at the current moment can be determined according to the position information of multiple marker points to be recognized at the current moment, and then the initial coordinate matrix can be translated according to the centroid coordinates to obtain a processed coordinate matrix.

[0068] More specifically, the centroid coordinates can be used as the origin of the coordinate system, and the three-dimensional coordinates of each marker point to be recognized can be updated to obtain a processed coordinate matrix. In other words, the coordinate of the i-th marker point to be recognized in the initial coordinate matrix is [x i , y i , z i , where i is a positive integer and 1≤i≤N. The coordinate of this marker point to be recognized in the processed coordinate matrix is [xi -x 0 , y i -y 0 , z i -z 0 , where the barycentric coordinates are [x 0 , y 0 , z 0 .

[0069] Adopting such a scheme can normalize the initial coordinate matrix, reduce the influence of the overall position of the object to be photographed on the category information, and is beneficial to improving the accuracy of the categories of the subsequent marker points to be recognized.

[0070] In the second specific example, the principal component analysis algorithm can also be used to process the initial coordinate matrix to obtain a rotation matrix, and then the rotation matrix is used to process the initial coordinate matrix to obtain a processed coordinate matrix.

[0071] Specifically, the principal component analysis algorithm can be used to process the initial coordinate matrix to obtain a rotation matrix, and then the product of the rotation matrix and the initial coordinate matrix is used as the processed coordinate matrix. Among them, the rotation matrix can be used to describe the mapping or transformation from the old three-dimensional rectangular coordinate system to the new three-dimensional rectangular coordinate system. The direction of the z-axis in the new three-dimensional rectangular coordinate system is the direction with the smallest variance of the coordinate matrix of multiple marker points to be recognized. By multiplying the rotation matrix with the initial coordinate matrix, the coordinates of each marker point to be recognized in the initial coordinate matrix under the old three-dimensional rectangular coordinate system can be transformed into the coordinates under the new three-dimensional rectangular coordinate system.

[0072] Adopting such a scheme can normalize the initial coordinate matrix, reduce the influence of the orientation of the object to be photographed on the category information, and is beneficial to improving the accuracy of the categories of the subsequent marker points to be recognized.

[0073] In the third specific example, the initial coordinate matrix can be processed according to the barycentric coordinates to obtain an intermediate coordinate matrix, and then the intermediate coordinate matrix can be processed according to the rotation matrix to obtain a processed coordinate matrix. Among them, the specific content of processing the initial coordinate matrix according to the barycentric coordinates and processing according to the rotation matrix can refer to the relevant descriptions above and will not be elaborated here.

[0074] It should be noted that the position information of the multiple marker points to be recognized in the embodiments of the present invention can be the coordinate matrix corresponding to the multiple marker points at the current moment. The coordinate matrix can be the above-mentioned initial coordinate matrix or the above-mentioned processed coordinate matrix, but is not limited thereto.

[0075] In the specific implementation of step S203, the position information of multiple to-be-recognized marker points is input into the classification model to obtain the category information of the multiple to-be-recognized marker points. More specifically, the preset category corresponding to the to-be-recognized marker point can be obtained, and thus the marker object corresponding to the to-be-recognized marker can be determined. It should be noted that if the preset category corresponding to the to-be-recognized marker point is the ghost point category, then the to-be-recognized marker has no corresponding marker object. If a certain preset category has no corresponding to-be-recognized marker point, then the marker object corresponding to this preset category is a lost marker object, and this preset category can be recorded as a missing point. Among them, the lost marker object may be lost during the process of taking the image or may be lost during the calculation of the coordinate matrix.

[0076] Among them, the classification model is pre-trained by using a training marker point set for a preset model. More content about training the preset model by using the training marker point set will be described below and will not be elaborated here.

[0077] Specifically, the classification model can be used to determine the category information of multiple to-be-recognized marker points according to the position information of the multiple to-be-recognized marker points. In other words, the input of the classification model is the position information of multiple to-be-recognized marker points at the current moment, and the output is the category information of multiple to-be-recognized marker points at the current moment.

[0078] More specifically, the input of the classification model can be the coordinate matrix corresponding to multiple to-be-recognized marker points at the current moment. The coordinate matrix can include the coordinates of each to-be-recognized marker point at the current moment. The coordinate matrix can be the initial coordinate matrix in step S202 or the processed coordinate matrix obtained by preprocessing the initial coordinate matrix in step S202. The embodiments of the present invention do not limit this.

[0079] The output of the classification model can be the probability matrix corresponding to multiple to-be-recognized marker points at the current moment. The probability matrix can include the probabilities of each to-be-recognized marker point belonging to each preset category. Among them, the preset category can be set in advance. In a specific example, multiple preset categories can include the ghost point category and multiple entity categories. Among them, there is a corresponding relationship between the entity category and the marker object set on the object to be photographed. For each marker object set on the object to be photographed, a unique entity category corresponding to it can be determined among the multiple entity categories.

[0080] In a specific example, the marker objects set on the object to be photographed are in one-to-one correspondence with the entity categories, that is, the number of marker objects set on the object to be photographed is the same as the number of preset entity categories.

[0081] In another specific example, the entity category corresponding to the marked object set on the object to be photographed is a subset of the preset entity category. In this case, for the probability matrix output by the classification model, the probability that each to-be-recognized annotation point belongs to the entity category without a corresponding marked object is removed to obtain an actual probability matrix, and the actual probability matrix is processed to obtain the category information of each to-be-recognized annotation point. Specifically, the actual probability matrix can be processed with reference to the method described later.

[0082] Refer to Figure 3 , Figure 3 is a schematic structural diagram of a classification model in an embodiment of the present invention. As Figure 3 shown, the classification model may include a first fully connected layer 31, a feature fusion module 32, and a first normalization module 33.

[0083] Specifically, the input of the first fully connected layer 31 may be the input of the classification model. The first fully connected layer 31 may be configured to calculate the feature information of each to-be-recognized marker point according to the position information of each to-be-recognized marker point, so as to obtain the feature information corresponding to multiple to-be-recognized marker points at the current moment. The feature information corresponding to multiple to-be-recognized marker points at the current moment includes the feature information of each to-be-recognized marker point.

[0084] More specifically, the input of the first fully connected layer 31 may be an N×3 coordinate matrix, and the output of the first fully connected layer 31 may be an N×K feature matrix, where K is a positive integer. More specifically, K may be a positive integer greater than 3. K is the feature dimension and K may be preset. In other words, the first fully connected layer 31 may convert the 1×3 coordinate information of each to-be-recognized marker point into 1×K feature information.

[0085] Further, the output of the first fully connected layer 31 may be connected to the input of the feature fusion module 32. The feature fusion module 32 may be configured to update the feature information of each to-be-recognized marker point according to the feature information of multiple to-be-recognized marker points in the neighborhood of each to-be-recognized marker point.

[0086] Specifically, the feature fusion module 32 may include a geometric information fusion sub-module (not shown in the figure). The geometric information fusion sub-module may be configured to calculate the first fusion feature information of each to-be-recognized marker point according to the feature information of the first adjacent marker point of each to-be-recognized marker point, and update the feature information of the to-be-recognized marker point with the first fusion feature information, that is, use the first fusion feature information of the to-be-recognized marker point as the feature information of the to-be-recognized marker point.

[0087] More specifically, the first adjacent marker points of each marker point to be recognized can be determined first. The first adjacent marker points are the geometric neighborhoods of the marker points to be recognized. In other words, the first adjacent marker points are the adjacent points in the Euclidean space determined according to the position information of each marker point to be recognized.

[0088] In a specific example, for each marker point to be recognized, the first preset number of marker points to be recognized with the shortest distance to the marker point to be recognized can be determined according to the position information of each marker point to be recognized, which are the first adjacent marker points of the marker point to be recognized, where the first preset number can be preset.

[0089] In another specific example, for each marker point to be recognized, a plurality of marker points to be recognized with a distance less than or equal to a preset distance threshold to the marker point to be recognized can be determined according to the position information of each marker point to be recognized, which are the first adjacent marker points of the marker point to be recognized, where the preset distance threshold can be preset.

[0090] Furthermore, the feature information of the first adjacent marker points (i.e., the feature information before fusion) can be fused to obtain the first fused feature information (i.e., the feature information after fusion). In a specific example, the geometric information fusion sub-module can include a pooling layer, and the feature information of the first adjacent marker points can be input into the pooling layer to obtain the first fused feature information. It should be noted that the specific process of the fusion process in the embodiments of the present invention is not limited, and various existing methods for fusing feature information can be used.

[0091] More specifically, the input of the geometric information fusion sub-module can be a feature matrix of N×K, the first preset number is J, the feature information of the first adjacent marker points can be a feature matrix of N×J×K, and the first fused feature information can be a feature matrix of N×K.

[0092] By setting the geometric information fusion sub-module, the feature information of the adjacent points of the marker points to be recognized in the Euclidean space can be fused, and the local feature information near each marker point to be recognized can be stably obtained, so that the perception field of the classification model can be expanded, which is beneficial to improving the accuracy of determining each marker point to be recognized.

[0093] Furthermore, the feature fusion module 32 can also include an edge convolution sub-module (not shown in the figure). The edge convolution sub-module can be used to calculate the second fused feature information of each marker point to be recognized according to the feature information of the second adjacent marker points of the marker point to be recognized, and update the feature information of the marker point to be recognized with the second fused feature information, that is, use the second fused feature information of the marker point to be recognized as the feature information of the marker point to be recognized.

[0094] Specifically, the second adjacent marker points of each marker point to be recognized can be determined first. Among them, the second preset number of marker points to be recognized with the smallest difference from the feature information of the marker point to be recognized can be determined according to the feature information of each marker point to be recognized, which are the second adjacent marker points of the marker point to be recognized. The second adjacent marker points are the feature neighborhood of the marker point to be recognized.

[0095] More specifically, the feature information is a feature vector. The second preset number of marker points to be recognized with the smallest distance of the feature vector from the marker point to be recognized can be used as the second adjacent marker points of the marker point to be recognized. In other words, the second adjacent marker points are the similar points in the feature space determined according to the feature information of each marker point to be recognized. Among them, the second preset number can be preset. The first preset number and the second preset number can be the same or different.

[0096] Furthermore, the feature information of the second adjacent marker points (i.e., the feature information before fusion) can be fused to obtain the second fused feature information (i.e., the feature information after fusion). In a specific example, the edge convolution sub-module may include a pooling layer. The feature information of the second adjacent marker points can be input into the pooling layer to obtain the second fused feature information. It should be noted that the embodiment of the present invention does not limit the specific process of the fusion process.

[0097] More specifically, the input of the edge convolution sub-module can be a feature matrix of N×K, the second preset number is Q, the feature information of the first adjacent marker points can be a feature matrix of N×Q×K, and the second fused feature information obtained after fusion can be a feature matrix of N×K.

[0098] By setting the edge convolution sub-module, the feature information of the similar points of the marker points to be recognized in the feature space can be fused. When the feature information of multiple marker points to be recognized changes, the feature neighborhood of each marker point to be recognized also changes, thereby greatly expanding the perception field of the classification model and facilitating improving the accuracy of determining each marker point to be recognized.

[0099] In the first specific example, the output of the first fully connected layer 31 can be connected to the input of the geometric information fusion sub-module, and the output of the geometric information sub-module can be connected to the input of the first normalization module 33. That is, the feature information fusion module 32 can only include the geometric information sub-module.

[0100] In the second specific example, the output of the first fully connected layer 31 can be connected to the input of the edge convolution sub-module, and the output of the edge convolution sub-module can be connected to the input of the first normalization module 33. That is, the feature information fusion module 32 can only include the edge convolution sub-module.

[0101] In the third specific example, the output of the first fully-connected layer 31 can be connected to the input of the geometric information fusion sub-module, the output of the geometric information fusion sub-module is connected to the input of the edge convolution sub-module, and the output of the edge convolution sub-module can be connected to the input of the first normalization module 33.

[0102] In the fourth specific example, the output of the first fully-connected layer 31 can be connected to the input of the edge convolution sub-module, the output of the edge convolution sub-module is connected to the input of the geometric information fusion sub-module, and the output of the geometric information fusion sub-module can be connected to the input of the first normalization module 33.

[0103] Furthermore, the output of the feature fusion module 32 is connected to the input of the first normalization module 33. The first normalization module 33 can be used to determine the category information of multiple to-be-recognized marker points according to the feature information of the multiple to-be-recognized marker points. More specifically, the first normalization module 33 can calculate a probability matrix based on the matrix output by the feature fusion module 32. In a specific example, the first normalization module 33 can obtain the probability matrix based on the Softmax function, but is not limited thereto.

[0104] Refer to Figure 4 , Figure 4 which is a schematic structural diagram of another classification model in the embodiments of the present invention. Figure 4 The classification model shown can include: a stretching processing module 41, a second fully-connected layer 42, and a second normalization module 43.

[0105] Specifically, the input of the stretching processing module 41 can be the input of the classification model. The stretching processing module 41 can be used to perform stretching processing on a multi-dimensional coordinate matrix to obtain a stretched coordinate matrix, where the stretched coordinate matrix is a one-dimensional matrix. More specifically, an N×3 coordinate matrix can be stretched into a 1×3N matrix. It should be noted that the embodiments of the present invention do not limit the specific manner of stretching.

[0106] Furthermore, the output of the stretching processing module 41 can be connected to the input of the second fully-connected layer 42. The second fully-connected layer 42 can be used to process the stretched coordinate matrix to obtain a one-dimensional output matrix, so as to obtain the feature information corresponding to the whole of multiple to-be-recognized marker points at the current moment. More specifically, the input of the second fully-connected layer 42 can be a 1×3N matrix, and the output can be 1×K', where K'>3N. The second fully-connected layer 42 can include multiple cascaded fully-connected layers, and there are residual connections between adjacent two fully-connected layers.

[0107] It should be noted that, compared with Figure 3Compared with the first fully connected layer 31 shown in Figure 4 , the first fully connected layer 31 calculates the feature information of each to-be-recognized marker point based on the position information of the to-be-recognized marker points, and the feature information of multiple to-be-recognized marker points is N-dimensional; while

[0108] the second fully connected layer 42 shown in

[0109] directly obtains the feature information of multiple to-be-recognized marker points at the current moment according to the stretched one-dimensional coordinate matrix, and the feature information of multiple to-be-recognized marker points is also one-dimensional. Figure 4 It should be noted that

[0110] the shown classification model is relatively lightweight, which is beneficial to improving the efficiency of determining the categories of to-be-recognized marker points.

[0111] From the above, the category information of multiple to-be-recognized marker points at the current moment can be obtained from the classification model. Specifically, the category information can be a probability matrix, and the probability matrix can include the probabilities of each to-be-recognized marker point belonging to each preset category, and the corresponding categories of each to-be-recognized marker point can be determined according to the probability matrix.

[0112] It should be noted that the predicted ghost point marker is not the to-be-recognized marker point with the highest probability of belonging to the ghost point category among multiple to-be-recognized marker points, but the to-be-recognized marker point with the largest probability of belonging to the ghost point category among the probabilities of each to-be-recognized marker point belonging to each preset category, and then this to-be-recognized marker point is used as the preset ghost point marker. Adopting such a scheme is beneficial to removing as many ghost points as possible.

[0113] Furthermore, the probabilities of the predicted ghost point marker belonging to each preset category and the probabilities of each to-be-recognized marker point belonging to the ghost point category can be excluded from the initial probability matrix to obtain a first probability matrix. It should be noted that compared with the initial probability matrix, the preset categories in the first probability matrix do not include the ghost point category, but only include entity categories.

[0114] In a specific example, the row vector corresponding to the predicted ghost point label and the column vector corresponding to the ghost point category in the initial probability matrix can be removed. More specifically, the initial probability matrix is an N×C matrix, where C is the number of preset categories, and the first probability matrix can be an M×(C - 1) matrix, where M ≤ N and M is a positive integer.

[0115] Further, determine the preset category corresponding to each to-be-identified marker point according to the first probability matrix. Specifically, for each to-be-identified marker point in the first probability matrix, the maximum value among the probabilities that the to-be-identified marker point belongs to each preset category in the first probability matrix can be determined, and the preset category corresponding to the maximum value is used as the preset category corresponding to the to-be-identified marker point.

[0116] In a second specific example, the first probability matrix can be further processed to obtain a second probability matrix, and the second probability matrix is a Doubly Stochastic Matrix. In a specific example, the Sinkhorn-Knopp normalization method can be used to process the first probability matrix to obtain the second probability matrix, but it is not limited thereto.

[0117] It should be noted that the number of rows and columns in the second probability matrix is the same as that in the first probability matrix. In other words, compared with the first probability matrix, the to-be-identified marker points and preset categories in the second probability matrix remain unchanged, only the corresponding probability values change. That is, the preset categories in the second probability matrix only include entity categories.

[0118] Further, the matching relationship between each to-be-identified marker point and the preset category can be determined according to the second probability matrix. Specifically, for each to-be-identified marker point in the second probability matrix, the maximum value among the probabilities that the to-be-identified marker point belongs to each preset category in the second probability matrix is determined, and the preset category corresponding to the maximum value is used as the preset category corresponding to the to-be-identified marker point.

[0119] Compared with the first probability matrix in the above text, since the second probability matrix is a doubly stochastic matrix, the values in the second probability matrix can also be used to indicate the probability that each preset category belongs to each to-be-identified marker point. Therefore, determining the matching relationship between each to-be-identified marker point and the preset category according to the second probability matrix can avoid the situation where multiple to-be-identified marker points belong to the same preset category. That is, the matching relationship between each to-be-identified marker point and the preset category can be determined. More specifically, for each to-be-identified marker point in the second probability matrix, a preset category can be matched, but for a preset category, it is not necessarily possible to match a to-be-identified marker point. For the preset category of the to-be-identified marker point that has no match, it can be determined that the marker object corresponding to the preset category is a lost marker object, that is, the preset category is a missing point.

[0120] In the third specific example, the cost matrix of the probability matrix can be determined, where the probability matrix can be the first probability matrix above, the second probability matrix above, or the initial probability matrix, and the embodiments of the present invention do not limit this. Specifically, the cost matrix can be calculated using the following formula:

[0121]

[0122] where, is the cost matrix, P is the probability matrix, and I is a matrix with all element values being 1. More specifically, the elements in the cost matrix correspond one-to-one with the elements in the probability matrix, and the value of each element in the cost matrix is the difference between 1 and the corresponding element in the probability matrix.

[0123] Further, according to the bipartite matching algorithm and the cost matrix, the preset category corresponding to each marker point to be recognized can be determined. By adopting such a solution, the matching relationship between each marker point to be recognized and the preset category can be determined. Among them, the bipartite matching algorithm can be the Hungarian algorithm, etc., but is not limited thereto.

[0124] By adopting such a solution, it is beneficial to avoid the situation that the marker points to be recognized belonging to ghost points are misjudged as entity marker points, thereby being beneficial to further improving the accuracy of the category information of each marker to be recognized.

[0125] It should be noted that if there is an entity category that does not have a corresponding marker point to be recognized, this entity category can be regarded as a missing point.

[0126] From the above, the preset category corresponding to each marker point to be recognized at the current moment can be obtained. More specifically, the entity category corresponding to each entity marker point at the current moment can be obtained, and the entity marker point is the image of the marker object set on the object to be photographed in the image.

[0127] Further, the action posture of the object to be photographed at the current moment can be determined according to the preset category corresponding to each marker point to be recognized at the current moment, and thus the action parameter sequence can be further determined according to the bone model of the object to be photographed. The action parameter sequence can be used to drive the virtual object to make the same action posture as the object to be photographed. It should be noted that the specific content of determining the action posture of the object to be photographed according to the category information of multiple markers to be recognized, and subsequently driving the virtual object according to the action posture of the object to be photographed can be various appropriate existing methods, and the embodiments of the present invention do not limit this.

[0128] In practical applications, if there are missing points in the obtained probability matrix, the action posture can be determined according to the entity category corresponding to the actually captured entity marker points. If there are ghost points in the probability matrix, the action posture is determined by the entity category corresponding to the entity marker points other than the ghost points.

[0129] Refer to Figure 5 , Figure 5 which is a schematic flowchart of a method for training the classification model in an embodiment of the present invention. Figure 5 The method shown may include steps S501 to S504. By executing steps S501 to S504, a preset model can be trained into a classification model. Steps S501 to S504 can be executed before step S203, or can be executed before step S201 or step S202. Figure 5 The method shown may include:

[0130] Step S501: Obtain the training marker point set;

[0131] Step S502: Input the position information of the multiple sample marker points into the preset model to obtain predicted category information;

[0132] Step S503: Calculate a loss value according to the predicted category information and the true category information;

[0133] Step S504: Update the preset model according to the loss value, and when a preset stop condition is satisfied, obtain the classification model.

[0134] In the specific implementation of step S501, the training marker point set may include the position information and true category information of multiple sample marker points of the subject in multiple action postures. Among them, the multiple to-be-identified marker points in the above text may be a subset of the multiple sample marker points.

[0135] Specifically, the training marker point set may include multiple groups of training data. Among them, each group of training data corresponds to an action posture, and each group of training data may include the position information and true category information of multiple sample marker points when the subject is in the action posture corresponding to this group of training data.

[0136] More specifically, each group of training data may include multiple data subgroups, and the data subgroups correspond one-to-one to the spatial state characteristics of the subject. Among them, the spatial state characteristics are used to describe the position where the whole subject is located and / or the orientation of the whole subject. Each data subgroup may include the position information and true category information of multiple sample marker points when the subject is in the corresponding action posture and has the corresponding spatial state characteristics. Among them, the position information may be obtained in advance, and the true category information may also be pre-annotated.

[0137] Furthermore, the training marker point set can be further enhanced, and the preset model can be trained using the enhanced training marker point set to improve the performance of the classification model. The enhancement processing may include one or more of the following: globally rotating and / or translating the position information of at least one set of sample marker points; adding information of ghost points; removing the information of at least one sample marker point in the training marker point set; adjusting the position information of at least one sample marker point in the training marker point set; adding the information of multiple sample marker points of the object to be photographed in a new pose; adding the information of multiple sample marker points of other objects other than the object to be photographed in the multiple action poses.

[0138] Specifically, the position information of multiple sample marker points in at least one data group can be globally rotated and / or translated to obtain new position information. Since in the same action pose, the orientation of the whole object to be photographed does not affect the category of the marker points, and in the same action pose, the position of the whole object to be photographed in space does not affect the category of the marker points. For example, when the object to be photographed is standing, lying, tilted, or facing different angles, the category of the marker points will not change. Therefore, adopting such a scheme is beneficial to making the classification model adapt to the situation where the same object to be photographed has different positions and / or orientations in the same action pose, which is conducive to enhancing the robustness of the classification model.

[0139] Furthermore, the information of ghost points can be added to one or more groups of training data. More specifically, multiple sample ghost marker points can be randomly generated, and the position information of the multiple sample ghost marker points can be random, and the true category information is the ghost point category.

[0140] Furthermore, the position information and true category information of at least one sample marker point can be removed to simulate the abnormal situation where some marker points may not be captured due to environmental occlusion and other reasons, thus there are missing points, which is conducive to improving the robustness of the classification model.

[0141] Furthermore, the position information of at least one sample marker point can be adjusted. By adopting such a scheme, the error of the sticker position caused by factors such as the state of the object to be photographed (for example, the muscle state of a real actor) during actual motion capture can be simulated, which is conducive to improving the robustness of the classification model.

[0142] Furthermore, the position information and category information of multiple sample marker points of the object to be photographed in a new pose can be added. Specifically, the new pose can be generated by processing the data of the existing action pose using the quaternion interpolation method. According to the new pose and the bone model of the object to be photographed, the position information of multiple sample marker points in the new pose can be obtained.

[0143] Furthermore, it is also possible to add information on multiple sample marker points of other objects other than the object to be photographed in multiple action postures. Specifically, a skeletal model of other objects can be selected, and the position information of multiple sample marker points can be calculated respectively according to the skeletal model of other objects and the data of the action postures. Adopting such a solution is beneficial to improving the robustness of the classification model for action capture of objects to be photographed in different forms.

[0144] In the specific implementation of step S502, the position information of multiple sample marker points can be input into a preset model to obtain predicted class information. Among them, the position information of multiple sample marker points can be a sample coordinate matrix, and the sample coordinate matrix can include the coordinates of each sample marker point. The predicted class information can be an initial predicted probability matrix, and the initial predicted probability matrix is a probability matrix output by the preset model. The initial predicted probability matrix can include the probability that each sample marker point belongs to each preset class. Regarding the structure of the preset model, reference can be made to the relevant description of the classification model in Figure 3 and Figure 4 above, and details will not be elaborated here.

[0145] It should be noted that before inputting the sample coordinate matrix into the preset model, the sample coordinate matrix can also be preprocessed to obtain a processed sample coordinate matrix. Preprocessing the sample coordinate matrix can include processing the sample coordinate matrix according to the centroid coordinates and processing the sample coordinate matrix according to the rotation matrix, but is not limited thereto. For more content on the sample coordinate matrix and preprocessing the sample coordinate matrix, reference can be made to Figure 1 the relevant description of step S202 in

[0146] In the specific implementation of step S503, the loss value can be calculated according to the predicted class information and the true class information. Specifically, the true class information can be a true class matrix, where the true class matrix can be a binary matrix. In other words, the true class matrix is used to indicate the true class of each sample marker point. More specifically, in the true class matrix, the probability that each sample marker point belongs to its true class is 1, and the probability of belonging to other preset classes is 0. Further, the loss value can be calculated according to a preset loss function, the predicted probability matrix, and the true class matrix.

[0147] In the first specific example, the cross-entropy method can be used to calculate the loss value according to the initial predicted probability matrix and the true class matrix, and the loss value can be used to indicate the difference between the initial predicted probability matrix and the true class matrix.

[0148] In a second specific example, multiple sample marker points may include zero or more ghost points and multiple sample entity marker points. One or more sample marker points may also be removed to simulate the situation where a marked object is lost. The probabilities that each sample marker point belongs to the ghost point category and the probabilities that the ghost points belong to each preset category may be excluded from the initial prediction probability matrix to obtain a first prediction probability matrix. Then, the first prediction probability matrix is processed to obtain a second prediction probability matrix, where the second prediction probability matrix is a doubly stochastic matrix. Further, the probabilities that each sample marker point belongs to the ghost point category and the probabilities that the ghost points belong to each preset category are excluded from the true category matrix to obtain a processed true category matrix. Then, the loss value is calculated based on the second probability matrix and the processed true category matrix. It should be noted that only entity categories are included in the first prediction probability matrix.

[0149] In a third specific example, the coordinates of the ghost points are excluded from the sample coordinate matrix to obtain a sample coordinate matrix after exclusion, where the sample coordinate matrix after exclusion only includes the coordinates of the sample entity marker points. Then, based on the sample coordinate matrix after exclusion and a third prediction probability matrix, a soft coordinate matrix is calculated. The soft coordinate matrix can be used to indicate the coordinates of the entity sample marker points corresponding to each entity category. It should be noted that the correspondence between the entity categories and the sample entity marker points in the soft coordinate matrix is obtained based on the prediction results of the prediction model.

[0150] Further, a predicted distance matrix can be calculated based on the soft coordinate matrix, where the predicted distance matrix includes the predicted distances between every two sample entity marker points.

[0151] Further, a loss value can be calculated based on the predicted distance matrix and the true distance matrix. The true distance matrix may include the true distances between every two marked objects set on the object to be photographed.

[0152] The parts of the object to be photographed usually have the property of a rigid body. The distances between the marked objects set on the same part usually do not change with the change of the action posture. When training a preset model using such a scheme, the rigid body distances between every two marked objects can be used as constraints for supervised learning, which can make the marked objects on the same rigid body in the prediction results conform to the restrictions of the rigid body properties, thus facilitating the improvement of the accuracy of the model.

[0153] In a non - restrictive example, a mask matrix can be used to process the predicted distance matrix and the ground - truth distance matrix respectively to obtain the processed predicted distance matrix and the processed ground - truth distance matrix. More specifically, for the missing labeled objects, the corresponding rows and columns of the labeled objects in the predicted distance matrix and the ground - truth distance matrix can be removed to obtain the processed predicted distance matrix and the ground - truth distance matrix, and then the loss value can be calculated based on the processed predicted distance matrix and the predicted ground - truth distance matrix. Adopting such a scheme can simulate the situation of missing labeled objects and is beneficial to improving the accuracy of the model.

[0154] In the specific implementation of step S504, the preset model can be updated according to the loss value. Specifically, the weights in the preset model can be adjusted according to the loss value. More specifically, each time the weights are adjusted, the weights in other modules except the normalization module in the preset model can be adjusted simultaneously according to the loss value. Other appropriate operations can also be performed on the preset model according to the loss value, and the embodiments of the present invention do not limit this. Among them, the method of adjusting the weights of the preset model can be various existing appropriate methods. For example, the weights of the preset model can be adjusted by using the gradient descent method. Before each update of the preset model, it can be first determined whether the preset stop condition is met. The preset stop condition can be various existing appropriate stop conditions used when training a neural network model. For example, the preset stop condition can be that the loss value converges to a preset loss threshold, or the number of epochs reaches a preset cycle threshold, where the preset loss threshold and the preset cycle threshold can be set in advance.

[0155] If the preset stop condition is met, the classification model can be obtained; if the preset stop condition is not met, the preset model can be updated and then return to step S502 until the preset stop condition is met. Regarding Figure 5 More content of the method shown can be referred to the relevant description above regarding Figures 1 to 4 and will not be elaborated here.

[0156] Refer to Figure 6 , Figure 6 is a schematic structural diagram of a device for determining the category of marker points for motion capture in an embodiment of the present invention. Figure 6 The device for determining the category of marker points for motion capture shown can include:

[0157] An image acquisition module 61, configured to acquire multiple images at the current moment, where each image includes multiple marker points to be recognized;

[0158] A position acquisition module 62, configured to determine the position information of the multiple marker points to be recognized at the current moment according to the multiple images;

[0159] The category prediction module 63 is configured to input the position information of the multiple to-be-recognized marker points into a classification model to obtain the category information of each to-be-recognized marker point. The classification model is obtained by training a preset model with training data, where the training data includes the position information and true category information of multiple sample marker points under multiple action postures.

[0160] In a specific implementation, the above-mentioned device for determining the category of marker points for motion capture may correspond to a chip with data processing capabilities in the terminal; or correspond to a chip module with data processing capabilities in the terminal, or correspond to the terminal.

[0161] Regarding Figure 6 For more content such as the working principle, working mode, and beneficial effects of the device for determining the category of marker points for motion capture shown, reference can be made to the relevant description above regarding Figures 1 to 5 and will not be elaborated here.

[0162] An embodiment of the present invention further provides a storage medium on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above-mentioned method for determining the category of marker points for motion capture. The storage medium may include ROM, RAM, a magnetic disk, or an optical disc, etc. The storage medium may also include a non-volatile memory or a non-transitory memory, etc.

[0163] An embodiment of the present invention further provides a terminal, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, it executes the steps of the above-mentioned method for determining the category of marker points for motion capture. The terminal includes, but is not limited to, terminal devices such as mobile phones, computers, and tablet computers.

[0164] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU for short), and this processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0165] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0166] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer program can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner.

[0167] In several embodiments provided in the present application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, in each embodiment of the present invention, the functional units can be integrated into a processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units. For example, for each device or product applied to or integrated into a chip, each module / unit included therein can be implemented in the form of hardware such as circuits, or at least some modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device or product applied to or integrated into a chip module, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device or product applied to or integrated into a terminal, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal, or at least some modules / units can be implemented in the form of software programs that run on the processor integrated inside the terminal, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits.

[0169] It should be understood that the term “and / or” in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character “ / ” in this article indicates that the associated objects before and after are in an “or” relationship.

[0170] In the embodiments of the present application, "a plurality of" means two or more.

[0171] In the embodiments of the present application, the descriptions such as first and second are only for illustration and distinguishing the described objects, without any order, nor do they represent special limitations on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.

[0172] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. A method for determining the category of marker points for motion capture, characterized in that, the method includes: Obtain multiple images at the current moment, where each image includes multiple marker points to be recognized; Determine the position information of the multiple marker points to be recognized at the current moment according to the multiple images; Input the position information of the multiple marker points to be recognized into a classification model to obtain the category information of the multiple marker points to be recognized, where the classification model is obtained by training a preset model with a training marker point set, and the training marker point set includes the position information and true category information of multiple sample marker points under multiple action postures; The category information of the multiple marker points to be recognized is an initial probability matrix, and the initial probability matrix includes the probability that each marker point to be recognized belongs to each preset category. The preset category includes a ghost point category. The ghost point is a marker point without a corresponding marker object. The method further includes: Determine a predicted ghost point marker, where the predicted ghost point marker is the marker point to be recognized with the highest probability of belonging to the ghost point category among the probabilities of belonging to each preset category; Remove the information of the predicted ghost point marker and the information of the ghost point category from the initial probability matrix to obtain a first probability matrix; Determine the preset category corresponding to each marker point to be recognized according to the first probability matrix; The determining the preset category corresponding to each marker point to be recognized according to the first probability matrix includes: Process the first probability matrix to obtain a second probability matrix, where the second probability matrix is a doubly stochastic matrix; Determine the preset category corresponding to each marker point to be recognized according to the second probability matrix; The determining the preset category corresponding to each marker point to be recognized according to the second probability matrix includes: In response to the preset categories corresponding to the multiple marker points to be recognized being a subset of all preset categories, remove the probability that each marker point to be recognized belongs to a missing point from the second probability matrix to obtain an actual probability matrix, and determine the preset category corresponding to each marker point to be recognized based on the actual probability matrix.

2. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, the method further includes: Determine the action posture of the object being photographed at the current moment according to the category information of the multiple marker points to be recognized.

3. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, The position information of the multiple marker points to be recognized is a coordinate matrix corresponding to the multiple marker points at the current moment. Determining the position information of the multiple marker points to be recognized according to the multiple images includes: Determine the initial coordinate matrix corresponding to the multiple marker points to be recognized according to the position of each marker point to be recognized in each image; Process the initial coordinate matrix according to the centroid coordinates and / or rotation matrix to obtain a processed coordinate matrix; Wherein, the centroid coordinates are the position of the centroid of the object being photographed at the current moment in the space coordinate system, and the rotation matrix is obtained by processing the initial coordinate matrix using the principal component analysis algorithm.

4. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, the classification model includes a first fully connected layer, a feature fusion module, and a first normalization module, wherein, the first fully connected layer is used to calculate the feature information of each marker point to be recognized according to the position information of the marker point to be recognized; the feature fusion module is used to update the feature information of each marker point to be recognized according to the feature information of multiple marker points to be recognized within the neighborhood of each marker point to be recognized; the first normalization module is used to determine the category information of the multiple marker points to be recognized according to the feature information of the multiple marker points to be recognized.

5. The method for determining the category of marker points for motion capture according to claim 4, characterized in that, the feature fusion module includes a geometric information fusion sub-module and / or an edge convolution sub-module: the geometric information fusion sub-module is used to update the feature information of each marker point to be recognized according to the feature information of the first adjacent marker points of each marker point to be recognized, wherein the first adjacent marker points are multiple marker points to be recognized within the geometric neighborhood of the marker point to be recognized; the edge convolution sub-module is used to update the feature information of each marker point to be recognized according to the feature information of the second adjacent marker points of each marker point to be recognized, wherein the second adjacent marker points are multiple marker points to be recognized within the feature neighborhood of the marker point to be recognized.

6. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, the position information of the multiple marker points to be recognized is a multi-dimensional coordinate matrix, the category information is a probability matrix, and the classification model includes: a stretching processing module, a second fully connected layer, and a second normalization module, wherein, the stretching processing module is used to perform stretching processing on the multi-dimensional coordinate matrix to obtain a stretched coordinate matrix, wherein the stretched coordinate matrix is a one-dimensional matrix; the second fully connected layer is used to process the stretched coordinate matrix to obtain a one-dimensional output matrix; the second normalization module is used to process the output matrix to obtain the probability matrix.

7. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, determining the preset category corresponding to each marker point to be recognized according to the second probability matrix further includes: for each marker point to be recognized in the second probability matrix, determining the maximum value among the probabilities that the marker point to be recognized belongs to each preset category in the second probability matrix, and taking the preset category corresponding to the maximum value as the preset category corresponding to the marker point to be recognized.

8. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, determining the preset category corresponding to each marker point to be recognized according to the first probability matrix includes: determining the cost matrix of the first probability matrix; according to the bipartite matching algorithm and the cost matrix, determining the preset category corresponding to each marker point to be recognized in the first probability matrix.

9. The method for determining the category of marker points for motion capture according to claim 1, characterized in that, The method for training the classification model includes: Obtain the training marker point set; Input the position information of the multiple sample marker points into the preset model to obtain predicted category information; Calculate a loss value according to the predicted category information and the true category information; Update the preset model according to the loss value, and when a preset stop condition is satisfied, obtain the classification model.

10. The method for determining the category of marker points for motion capture according to claim 9, wherein, the multiple sample marker points include ghost points and multiple sample entity marker points, the predicted category information is an initial predicted probability matrix, the initial predicted probability matrix includes the probability that each sample marker point belongs to each preset category, the preset categories include the ghost point category, and calculating the loss value according to the predicted category information and the true category information includes: Exclude the probabilities that the ghost points belong to each preset category and the probabilities that each sample marker point belongs to the ghost point category from the initial predicted probability matrix to obtain a first predicted probability matrix; Calculate the loss value according to the first predicted probability matrix and the true category information.

11. The method for determining the category of marker points for motion capture according to claim 10, wherein, calculating the loss value according to the first predicted probability matrix and the true category information includes: Process the first predicted probability matrix to obtain a second predicted probability matrix, and the second predicted probability matrix is a doubly stochastic matrix; Calculate the loss value according to the second predicted probability matrix and the true category information.

12. The method for determining the category of marker points for motion capture according to claim 10, wherein, the position information of the multiple sample marker points is a sample coordinate matrix, and calculating the loss value according to the first predicted probability matrix and the true category information includes: Exclude the coordinates of the ghost points from the sample coordinate matrix to obtain an excluded sample coordinate matrix; Calculate a predicted distance matrix according to the excluded sample coordinate matrix and the first predicted probability matrix, and the predicted distance matrix includes the predicted distances between every two sample entity marker points; Calculate the loss value according to the predicted distance matrix and the true distance matrix.

13. The method for determining the category of marker points for motion capture according to claim 9, wherein, before inputting the position information of the multiple sample marker points into the preset model, the method further includes: Perform enhancement processing on the training marker point set, and the enhancement processing includes one or more of the following: Perform global rotation and / or global translation processing on the position information of at least one group of sample marker points; Add information of ghost points; Remove the information of at least one sample marker point from the training marker point set; Adjust the position information of at least one sample marker point in the training marker point set; Add information of multiple sample marker points of the subject in a new pose; Add information of multiple sample marker points of other objects other than the subject in the multiple action poses.

14. An apparatus for determining the category of marker points for motion capture It is characterized in that the device includes an image acquisition module for acquiring multiple images at the current moment, where each image includes multiple to-be-recognized marker points; a position acquisition module for determining the position information of the multiple to-be-recognized marker points at the current moment according to the multiple images; a category prediction module for inputting the position information of the multiple to-be-recognized marker points into a classification model to obtain the category information of the multiple to-be-recognized marker points; wherein, the classification model is obtained by training a preset model with training data, and the training data includes the position information and true category information of multiple sample marker points under multiple action postures; the category information of the multiple to-be-recognized marker points is an initial probability matrix, the initial probability matrix includes the probability that each to-be-recognized marker point belongs to each preset category, the preset categories include the ghost point category, the ghost point is a marker point without a corresponding marker object, and the device further includes determining a predicted ghost point marker, where the predicted ghost point marker is the to-be-recognized marker point with the highest probability of belonging to the ghost point category among the probabilities of belonging to each preset category; removing the information of the predicted ghost point marker and the information of the ghost point category from the initial probability matrix to obtain a first probability matrix; determining the preset category corresponding to each to-be-recognized marker point according to the first probability matrix; the determining the preset category corresponding to each to-be-recognized marker point according to the first probability matrix includes processing the first probability matrix to obtain a second probability matrix, where the second probability matrix is a doubly stochastic matrix; determining the preset category corresponding to each to-be-recognized marker point according to the second probability matrix; the determining the preset category corresponding to each to-be-recognized marker point according to the second probability matrix includes in response to the preset categories corresponding to the multiple to-be-recognized marker points being a subset of all preset categories, removing the probability that each to-be-recognized marker point belongs to a missing point from the second probability matrix to obtain an actual probability matrix, and determining the preset category corresponding to each to-be-recognized marker point based on the actual probability matrix.

15. A storage medium, on which a computer program is stored, it is characterized in that when the computer program is run by a processor, it executes the steps of the method for determining the category of marker points for action capture according to any one of claims 1 to 13.

16. A terminal, including a memory and a processor, and a computer program that can run on the processor is stored on the memory, it is characterized in that when the processor runs the computer program, it executes the steps of the method for determining the category of marker points for action capture according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Rigid body motion capturing method and system based on mark point

    CN106600627A

  • Human posture recognition device and method based on full convolutional neural network

    CN108898063A

  • Recognition method and device for key points in image, electronic equipment and storage medium

    CN111753796A