Children video PK audience interactive guessing reward system and method
By normalizing the frame rate of children's videos and matching keyframes, combined with the space-time network graph scoring model, the problem of action recognition error in the existing technology is solved, and higher user experience and action recognition accuracy are achieved.
Patent Information
- Application Number
- CN202510395056.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, it is difficult to improve user experience while ensuring the accuracy of action recognition in children's video interactive entertainment, especially in the case of inconsistent video frames, action recognition is prone to errors.
The video processing module is used to normalize the frame rate of user videos and standard videos, and the keyframe images are extracted using the human key point recognition model, and keyframe matching is performed through dynamic time regularization algorithms, and a spatiotemporal network diagram is built to score the smoothness of the action, and finally a comprehensive score is generated through the comprehensive scoring module.
It realizes the improvement of user experience while ensuring the accuracy of action recognition. Through the combination of dynamic time regularization and space-time network graph scoring model, the accuracy of action recognition and rating is improved.
Smart Images

Figure CN120220032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and more specifically, it relates to a children's video PK audience interactive guessing reward system and method. Background Art
[0002] With the rapid development of technology, the field of children's video interactive entertainment has gradually emerged. In particular, the guessing reward system combined with audience participation has brought a new immersive experience to children. However, in the prior art, in order to achieve accurate action recognition and interactive feedback, usually two main methods are adopted: one is to require users to wear sensor devices to capture action data. Although this method can provide high-precision action recognition, the wearing of sensors limits the freedom of movement of users and affects the performance of natural stretching; the other is to collect the action videos of users and perform matching analysis by comparing each frame with a standard comparison video frame by frame. However, in practical applications, there may be problems with inconsistent video frames. For example, when the action frequency of the user is slightly faster or slower, even if the overall action completion is very high, there may be errors in action recognition due to the lack of semantic understanding of the action.
[0003] Therefore, how to improve the user experience while ensuring the recognition accuracy has become an urgent technical problem to be solved currently. Summary of the Invention
[0004] The present invention provides a children's video PK audience interactive guessing reward system and method to solve the technical problems in the above background art.
[0005] The present invention provides a children's video PK audience interactive guessing reward system, including: A video processing module 101, which is used to perform frame rate normalization processing on the user video according to the standard video, and at the same time perform frame splitting and grayscale processing on the standard video and the user video respectively to obtain a standard image set and a user image set; A key frame extraction module 102, which is used to identify the skeletal key points of each frame image in the standard image set and the user image set through a human key point recognition model, and extract key frame images by calculating the change amount of the skeletal key points between adjacent frames to obtain a first standard key frame set and a first user key frame set; A key frame matching module 103, which is used to perform key frame matching on the first standard key frame set and the first user key frame set through a dynamic time warping algorithm to obtain a second standard key frame set and a second user key frame set with the same number of key frame images; An action completion degree scoring module 104, which is used to calculate an action completion degree score according to the second standard key frame set and the second user key frame set; The spatio-temporal network graph construction module 105 is used to construct a spatio-temporal network graph according to the second set of user key frames; The spatio-temporal network graph is composed of the skeletal key points in all the key frame images in the second set of user key frames; The action fluency scoring module 106 is used to input the spatio-temporal network graph into a scoring model and output an action fluency score.
[0006] Furthermore, the frame rate normalization process includes the following steps: Step S201, if it is determined that the frame rate of the user video is less than the frame rate of the standard video, then calculate the number of frames to be inserted into the user video and the positions according to the standard video, so that the frame rate of the user video is the same as the frame rate of the standard video; The specific calculation formula is as follows: ; ; where and respectively represent the frame rate and duration of the standard video, and respectively represent the frame rate and duration of the user video, represents the number of frames to be inserted into the user video, represents the positions to be inserted into the user video; Step S202, if it is determined that the frame rate of the user video is greater than the frame rate of the standard video, then calculate the number of frames to be deleted from the user video and the positions according to the standard video, so that the frame rate of the user video is the same as the frame rate of the standard video; The specific calculation formula is as follows: ; ; where represents the number of frames to be deleted from the user video, represents the positions to be deleted from the user video.
[0007] Furthermore, the human key point recognition model is the OpenPose model.
[0008] Furthermore, the action completion degree score is calculated according to the second set of standard key frames and the second set of user key frames The calculation formula is as follows: ; where M represents the number of key frame images in the second set of standard key frames or the second set of user key frames, N represents the number of skeletal key points in any one key frame image, and respectively represent the Euclidean distance and cosine similarity between the nth skeletal key point in the mth key frame image of the second standard key frame set and the second user key frame set, and respectively represent the custom first weight coefficient and second weight coefficient with the sum value of 1.
[0009] Further, the number of subgraphs of the spatio-temporal network graph is the same as the number of key frame images in the second user key frame set; A node of a subgraph corresponds to a skeletal key point of a key frame image; A node of a subgraph is represented by an initial vector, and the dimension values of the initial vector include: the abscissa value, the ordinate value, and the gray value in the key frame image where the node-corresponding skeletal key point is located; The edges between nodes include: constructing edges between nodes corresponding to the same skeletal key point in adjacent key frame images; constructing edges between nodes corresponding to adjacent skeletal key points in the same key frame image.
[0010] Further, the scoring model includes M hidden layers, a non-linear transformation layer, and a classifier, where the number M of hidden layers is the same as the number of subgraphs of the spatio-temporal network graph; The mth hidden layer inputs the mth subgraph of the spatio-temporal graph network and outputs an update matrix, where 1 ≤ m ≤ M, and the update matrix is stacked by update vectors of N nodes, where the number N of row vectors of the update matrix is the same as the number of skeletal key points in a key frame image; The non-linear transformation layer inputs the update matrix output by the Mth hidden layer and outputs a feature vector; The classifier inputs the feature vector and outputs an action fluency score; The sample labels of the training samples for training the scoring model are manually labeled.
[0011] Further, the calculation formula of the mth hidden layer includes: ; where 1 ≤ n ≤ N, represents the update vector of the nth node of the update matrix output by the mth hidden layer, represents the initial vector of the nth node of the (m - 1)th subgraph input to the (m - 1)th hidden layer, and respectively represent the initial vectors of the nth node and the kth node of the mth subgraph input to the mth hidden layer, and respectively represent the node set and the number of nodes having an edge with the nth node of the mth subgraph, and respectively represent the time control vector and the space control vector corresponding to the nth node of the mth sub-graph, 、 and respectively represent the first weight matrix, the second weight matrix and the third weight matrix of the mth hidden layer, represents the sigmoid activation function, represents element-wise multiplication; ; where represents the fourth weight matrix of the mth hidden layer, represents the first bias vector of the mth hidden layer, GTU represents the gated activation function, GELU represents the GELU activation function, and Concat represents the concatenation function; ; where represents the fifth weight matrix of the mth hidden layer, represents the second bias vector of the mth hidden layer.
[0012] Furthermore, the calculation formula of the non-linear conversion layer is as follows: ; where Vector represents the feature vector output by the non-linear conversion layer, represents the updated matrix output by the Mth hidden layer, W and b respectively represent the weight vector and the bias vector of the non-linear conversion layer, and Swish represents the Swish activation function.
[0013] Furthermore, a children's video PK audience interactive guessing reward system provided by the present invention further includes: a comprehensive scoring module 107, which is used to perform weighted summation on the action completion score and the action fluency score to obtain a comprehensive score; The weight coefficients corresponding to the action completion score and the action fluency score are both user-defined parameters; An action feedback module 108, which is used to traverse and calculate the angle deviation values between all the connecting lines of the skeletal key points in the key frame images of the second standard key frame set and the second user key frame set, and if it is determined that the angle deviation value is greater than or equal to the angle deviation threshold, then label the connecting lines of the skeletal key points in the key frame image and return them to the user.
[0014] The present invention provides a children's video PK audience interactive guessing reward method, including the following steps: Step S301, perform frame rate normalization processing on the user video according to the standard video, and at the same time perform frame splitting and grayscale processing on the standard video and the user video respectively to obtain a standard image set and a user image set; Step S302: Identify the skeletal key points of each sub-frame image in the standard image set and the user image set through the human key point recognition model, and extract the key frame images by calculating the change amount of the skeletal key points between adjacent frames to obtain the first standard key frame set and the first user key frame set; Step S303: Perform key frame matching on the first standard key frame set and the first user key frame set through the dynamic time warping algorithm to obtain the second standard key frame set and the second user key frame set with the same number of key frame images; Step S304: Calculate the action completion degree score based on the second standard key frame set and the second user key frame set; Step S305: Construct a spatio-temporal network graph according to the second user key frame set; Step S306: Input the spatio-temporal network graph into the scoring model and output the action fluency score; Step S307: Perform weighted summation on the action completion degree score and the action fluency score to obtain the comprehensive score; Step S308: Traverse and calculate the angle deviation values between all the connected lines of the skeletal key points in the key frame images of the second standard key frame set and the second user key frame set. If the angle deviation value is greater than or equal to the angle deviation threshold, then label the connected lines of the skeletal key points in the key frame image and return it to the user.
[0015] The beneficial effects of the present invention are as follows: The present invention realizes the key frame matching of the user video and the standard video through dynamic time warping, thereby ensuring the consistency of the video frames, and adding semantic understanding of the user video through the scoring model, thereby improving the accuracy of action recognition and action scoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic diagram of a children's video PK audience interactive guessing reward system of the present invention; Figure 2 is a flowchart of the frame rate normalization process of the present invention; Figure 3 is a flowchart of a children's video PK audience interactive guessing reward method of the present invention; Figure 4 is a schematic diagram of the human key point recognition of the present invention.
[0017] In the figure: video processing module 101, key frame extraction module 102, key frame matching module 103, action completion degree scoring module 104, spatio-temporal network graph construction module 105, action fluency scoring module 106, comprehensive scoring module 107, action feedback module 108. DETAILED DESCRIPTION OF THE INVENTION
[0018] Reference will now be made to example embodiments to discuss the subject matter described herein. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.
[0019] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by those of ordinary skill in the art to which the present invention pertains. The "first", "second" and similar terms used in one or more embodiments of the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0020] As Figures 1 to 4 shown, a children's video PK audience interactive guessing reward system includes: A video processing module 101, which is used to perform frame rate normalization processing on the user video according to the standard video, and at the same time perform frame splitting and grayscale processing on the standard video and the user video respectively to obtain a standard image set and a user image set; A key frame extraction module 102, which is used to identify the skeletal key points of each frame-split image in the standard image set and the user image set through a human key point recognition model, and extract key frame images by calculating the change amount of the skeletal key points between adjacent frames to obtain a first standard key frame set and a first user key frame set; A key frame matching module 103, which is used to perform key frame matching on the first standard key frame set and the first user key frame set through a dynamic time warping algorithm to obtain a second standard key frame set and a second user key frame set with the same number of key frame images; An action completion degree scoring module 104, which is used to calculate the action completion degree score according to the second standard key frame set and the second user key frame set; A spatio-temporal network graph construction module 105, which is used to construct a spatio-temporal network graph according to the second user key frame set; The spatio-temporal network diagram is composed of the skeletal key points in all key frame images in the second user key frame set; The action fluency scoring module 106 is configured to input the spatio-temporal network diagram into a scoring model and output an action fluency score.
[0021] In an embodiment of the present invention, as Figure 2 shown, the frame rate normalization process includes the following steps: Step S201, determine that the frame rate of the user video is less than the frame rate of the standard video, and then calculate the number of frames to be inserted into the user video and the positions according to the standard video, so that the frame rate of the user video is consistent with the frame rate of the standard video; The specific calculation formula is as follows: ; ; where and respectively represent the frame rate and duration of the standard video, and respectively represent the frame rate and duration of the user video, represents the number of frames to be inserted into the user video, represents the positions to be inserted into the user video; Step S202, determine that the frame rate of the user video is greater than the frame rate of the standard video, and then calculate the number of frames to be deleted from the user video and the positions according to the standard video, so that the frame rate of the user video is consistent with the frame rate of the standard video; The specific calculation formula is as follows: ; ; where represents the number of frames to be deleted from the user video, represents the positions to be deleted from the user video.
[0022] It should be noted that frames adjacent to the insertion positions can be selected for insertion.
[0023] In an embodiment of the present invention, the calculation formula for grayscale processing is as follows: ; where represents the grayscale value, , and respectively represent the red channel value, green channel value and blue channel value.
[0024] In an embodiment of the present invention, the human key point recognition model is the OpenPose model, or it can also be AlphaPose, BlazePose, etc.
[0025] It should be noted that, first, the coordinate sequences of the skeletal key points of each key frame image are extracted from the first standard key frame set and the first user key frame set, and then a similarity matrix is obtained by calculating the Euclidean distance. The rows of the similarity matrix represent the number of key frame images in the first standard key frame set, and the columns of the similarity matrix represent the number of key frame images in the first user key frame set. Then, the dynamic time warping algorithm is used to find the optimal path from the similarity matrix.
[0026] In an embodiment of the present invention, an action completion degree score is calculated according to the second standard key frame set and the second user key frame set The calculation formula is as follows: ; where M represents the number of key frame images in the second standard key frame set or the second user key frame set, N represents the number of skeletal key points in any key frame image, and respectively represent the Euclidean distance and cosine similarity between the nth skeletal key point in the mth key frame image of the second standard key frame set and the second user key frame set, and respectively represent a custom first weight coefficient and a second weight coefficient with a sum value of 1. Preferably, is set to 0.4, is set to 0.6.
[0027] In an embodiment of the present invention, the number of subgraphs of the spatio-temporal network graph is the same as the number of key frame images in the second user key frame set; A node of a subgraph corresponds to a skeletal key point of a key frame image; A node of a subgraph is represented by an initial vector, and the dimension values of the initial vector include: the abscissa value, the ordinate value, and the gray value in the key frame image where the node corresponds to the skeletal key point; The edges between nodes include: edges are constructed between nodes corresponding to the same skeletal key point in adjacent key frame images; edges are constructed between nodes corresponding to adjacent skeletal key points in the same key frame image.
[0028] In an embodiment of the present invention, the scoring model includes M hidden layers, a non-linear transformation layer, and a classifier, where the number M of hidden layers is the same as the number of subgraphs of the spatio-temporal network graph; The mth hidden layer inputs the mth subgraph of the spatio-temporal graph network and outputs an update matrix, where 1 ≤ m ≤ M. The update matrix is stacked by update vectors of N nodes, where the number N of row vectors of the update matrix is the same as the number of skeletal key points in a key frame image; The update matrix of the input of the non - linear conversion layer to the output of the M - th hidden layer outputs the feature vector; The classifier inputs the feature vector and outputs the action smoothness score; The sample labels of the training samples for training the scoring model are manually labeled.
[0029] In an embodiment of the present invention, the calculation formula of the m - th hidden layer includes: ; where 1 ≤ n ≤ N, represents the update vector of the n - th node of the update matrix output by the m - th hidden layer, represents the initial vector of the n - th node of the (m - 1) - th sub - graph input to the (m - 1) - th hidden layer, and respectively represent the initial vectors of the n - th node and the k - th node of the m - th sub - graph input to the m - th hidden layer, and respectively represent the set of nodes and the number of nodes having an edge with the n - th node of the m - th sub - graph, and respectively represent the time control vector and the space control vector corresponding to the n - th node of the m - th sub - graph, 、 and respectively represent the first weight matrix, the second weight matrix and the third weight matrix of the m - th hidden layer, represents the sigmoid activation function, represents element - wise multiplication; ; where represents the fourth weight matrix of the m - th hidden layer, represents the first bias vector of the m - th hidden layer, GTU represents the gated activation function, GELU represents the GELU activation function, and Concat represents the concatenation function; ; where represents the fifth weight matrix of the m - th hidden layer, represents the second bias vector of the m - th hidden layer.
[0030] It should be noted that the present invention uses Markov's idea to design the time control vector and the space control vector, wherein the time control vector is used to control the influence of the vector change at the previous moment on the vector change at the next moment, and the space control vector is used to control the influence of the vector associated at the same moment on the vector change at the next moment, thereby realizing the association analysis in the time dimension and the space dimension. In addition, the above weight parameters and bias parameters are all learnable hyperparameters. For example, if the size of the initial vector is 1×3, then , and Can be designed as a 3×16 matrix, then and The size needs to be designed to 1×16 at the same time. and It needs to be designed as a 3×16 matrix. and It needs to be designed as a vector of size 1×16.
[0031] In one embodiment of the present invention, the calculation formula of the nonlinear conversion layer is as follows: ; Where Vector represents the feature vector output by the nonlinear transformation layer, represents the update matrix of the output of the Mth hidden layer, W and b represent the weight vector and bias vector of the nonlinear transformation layer respectively, and Swish represents the Swish activation function.
[0032] It should be noted that the weight vector and bias vector in the nonlinear transformation layer are also learnable hyperparameters, and the mean square error between the value output by the scoring model during training and the true value of the sample label of the training sample is specified as the loss function. The parameters in the scoring model are updated by the gradient descent algorithm, such as Adam, RMSprop, etc., which will not be elaborated here.
[0033] In one embodiment of the present invention, the children's video PK audience interactive guessing reward system provided by the present invention further includes: a comprehensive scoring module 107, which is used to perform a weighted summation of the action completion score and the action fluency score to obtain a comprehensive score; The weight coefficients corresponding to the action completion score and the action fluency score are both custom parameters. Preferably, the weight coefficient corresponding to the action completion score is set to 0.6, and the weight coefficient corresponding to the action fluency score is set to 0.4; The action feedback module 108 is used to traverse and calculate the angular deviation values between all the connected lines of the skeletal key points in the key frame images of the second standard key frame set and the second user key frame set. If the angular deviation value is greater than or equal to the angular deviation threshold, the connected lines of the skeletal key points in the key frame image are marked and returned to the user. As Figure 4 shown, it is a schematic diagram of identifying the connected lines of skeletal key points through the OpenPose model.
[0034] It should be noted that in the present invention, the key frames of the user video and the standard video are matched through dynamic time warping, so as to ensure the consistency of video frames, which is convenient for calculating the action completion score. The action completion score is mainly used to reflect the degree to which the user specifically completes compared with the standard video. The scoring model incorporates semantic understanding of the user video. Even if the action frequency of the user is slightly faster or slower, the scoring model can still score correctly. In addition, a penalty term for the action completion duration score can be added based on the comprehensive score, which will not be elaborated here.
[0035] In an embodiment of the present invention, as Figure 3 shown, a method for rewarding children's video PK audience interaction guessing includes the following steps: Step S301, perform frame rate normalization processing on the user video according to the standard video, and at the same time perform frame splitting and grayscale processing on the standard video and the user video respectively to obtain a standard image set and a user image set; Step S302, identify the skeletal key points of each frame image in the standard image set and the user image set through a human key point recognition model, and extract key frame images by calculating the change amount of the skeletal key points between adjacent frames to obtain a first standard key frame set and a first user key frame set; Step S303, perform key frame matching on the first standard key frame set and the first user key frame set through the dynamic time warping algorithm to obtain a second standard key frame set and a second user key frame set with the same number of key frame images; Step S304, calculate the action completion score according to the second standard key frame set and the second user key frame set; Step S305, construct a spatio-temporal network graph according to the second user key frame set; Step S306, input the spatio-temporal network graph into the scoring model to output the action fluency score; Step S307, perform weighted summation on the action completion score and the action fluency score to obtain a comprehensive score; Step S308, traverse and calculate the angular deviation values between the connecting lines of all skeletal key points in the key frame images of the second standard key frame set and the second user key frame set. If it is determined that the angular deviation value is greater than or equal to the angular deviation threshold, then label the connecting lines of the skeletal key points in this key frame image and return them to the user.
[0036] The above has described the embodiments of this embodiment, but this embodiment is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this embodiment.
Claims
1. A children's video PK audience interactive guessing reward system, characterized in that: include: The video processing module 101 is used to perform frame rate normalization processing on the user video according to the standard video, and simultaneously perform frame division and grayscale processing on the standard video and the user video to obtain a standard image set and a user image set respectively; A key frame extraction module 102 is used to identify the skeleton key points of each frame image in the standard image set and the user image set through a human body key point recognition model, and extract the key frame images by calculating the variation of the skeleton key points between adjacent frames to obtain a first standard key frame set and a first user key frame set; A key frame matching module 103, which is used to perform key frame matching on the first standard key frame set and the first user key frame set by using a dynamic time warping algorithm to obtain a second standard key frame set and a second user key frame set with the same number of key frame images; An action completion score module 104, which is used to calculate the action completion score according to the second standard key frame set and the second user key frame set; A spatiotemporal network graph construction module 105, which is used to construct a spatiotemporal network graph according to the second user key frame set; The spatiotemporal network graph is composed of the skeleton key points in all key frame images in the second user key frame set; The action fluency scoring module 106 is used to input the spatiotemporal network graph into the scoring model and output the action fluency score.
2. A children's video PK audience interactive guessing reward system according to claim 1, characterized in that: The frame rate normalization process includes the following steps: Step S201, if it is determined that the frame rate of the user video is lower than the frame rate of the standard video, the number of frames and the position of the inserted user video are calculated according to the standard video so that the frame rate of the user video is consistent with the frame rate of the standard video; The specific calculation formula is as follows: ; ; in and Respectively represent the frame rate and duration of the standard video, and Respectively represent the frame rate and duration of the user’s video, Indicates the number of frames inserted into the user video. Indicates the location where the user video is inserted; Step S202: if it is determined that the frame rate of the user video is greater than the frame rate of the standard video, the number of frames and the position of the user video to be deleted are calculated according to the standard video so that the frame rate of the user video is consistent with the frame rate of the standard video; The specific calculation formula is as follows: ; ; in Indicates the number of frames to delete from the user's video. Indicates the location where the user's video was deleted.
3. A children's video PK audience interactive guessing reward system according to claim 1, characterized in that: The human key point recognition model is the OpenPose model.
4. A children's video PK audience interactive guessing reward system according to claim 1, characterized in that: The action completion score is calculated based on the second standard key frame set and the second user key frame set. The calculation formula is as follows: ; Where M represents the number of key frame images in the second standard key frame set or the second user key frame set, and N represents the number of skeleton key points in any key frame image. and represent the Euclidean distance and cosine similarity between the nth skeleton key point in the mth key frame image of the second standard key frame set and the second user key frame set, respectively. and They respectively represent the customized first weight coefficient and second weight coefficient whose sum is 1.
5. The children's video PK audience interactive guessing reward system according to claim 1, characterized in that: The number of subgraphs of the spatiotemporal network graph is the same as the number of key frame images in the second user key frame set; A node of a subgraph corresponds to a skeleton keypoint of a keyframe image; A node of a subgraph is represented by an initial vector, and the dimension values of the initial vector include: the horizontal coordinate value, the vertical coordinate value and the gray value in the key frame image where the skeleton key point corresponding to the node is located; The edges between nodes include: edges constructed between nodes corresponding to the same skeleton key point in adjacent key frame images; edges constructed between nodes corresponding to adjacent skeleton key points in the same key frame image.
6. A children's video PK audience interactive guessing reward system according to claim 5, characterized in that: The scoring model includes M hidden layers, nonlinear transformation layers and classifiers, where the number of hidden layers M is the same as the number of subgraphs of the spatiotemporal network graph; The mth hidden layer inputs the mth subgraph of the spatiotemporal graph network and outputs an update matrix, where 1≤m≤M. The update matrix is formed by stacking the update vectors of N nodes, where the number of row vectors N in the update matrix is the same as the number of skeletal key points in a keyframe image; The nonlinear transformation layer inputs the update matrix output by the Mth hidden layer and outputs the feature vector; The classifier inputs the feature vector and outputs the movement fluency score; The sample labels of the training samples used to train the scoring model are manually annotated.
7. A children's video PK audience interactive guessing reward system according to claim 6, characterized in that: The calculation formula for the mth hidden layer includes: ; Where 1≤n≤N, represents the update vector of the nth node of the update matrix for the output of the mth hidden layer, represents the initial vector of the nth node of the m-1th subgraph input to the m-1th hidden layer, and They represent the initial vectors of the nth node and the kth node of the mth subgraph input to the mth hidden layer, respectively. and They represent the node set and number of nodes that have edges with the nth node of the mth subgraph, and denote the time control vector and space control vector corresponding to the nth node of the mth subgraph, respectively. , and denote the first weight matrix, the second weight matrix and the third weight matrix of the mth hidden layer respectively, represents the sigmoid activation function, represents point-wise multiplication; ; in represents the fourth weight matrix of the mth hidden layer, represents the first bias vector of the mth hidden layer, GTU represents the gated activation function, GELU represents the GELU activation function, and Concat represents the concatenation function; ; in represents the fifth weight matrix of the mth hidden layer, Represents the second bias vector of the mth hidden layer.
8. The children's video PK audience interactive guessing reward system according to claim 6, characterized in that: The calculation formula of the nonlinear transformation layer is as follows: ; Where Vector represents the feature vector output by the nonlinear transformation layer, represents the update matrix of the output of the Mth hidden layer, W and b represent the weight vector and bias vector of the nonlinear transformation layer respectively, and Swish represents the Swish activation function.
9. The children's video PK audience interactive guessing reward system according to claim 1, characterized in that: The children's video PK audience interactive guessing reward system provided by the present invention also includes: a comprehensive scoring module 107, which is used to perform weighted summation of the action completion score and the action fluency score to obtain a comprehensive score; The weight coefficients corresponding to the action completion score and action fluency score are both custom parameters; The action feedback module 108 is used to traverse and calculate the angle deviation values between all the skeletal key point lines in the key frame images of the second standard key frame set and the second user key frame set, and if it is determined that the angle deviation value is greater than or equal to the angle deviation threshold, the skeletal key point lines in the key frame image are marked and returned to the user.
10. A children's video PK audience interactive guessing reward method, characterized in that: Executing a children's video PK audience interactive guessing reward system as described in any one of claims 1 to 9 comprises the following steps: Step S301, performing frame rate normalization processing on the user video according to the standard video, and performing frame division and grayscale processing on the standard video and the user video to obtain a standard image set and a user image set respectively; Step S302, identifying the skeleton key points of each frame image in the standard image set and the user image set by using a human body key point recognition model, and extracting key frame images by calculating the variation of the skeleton key points between adjacent frames to obtain a first standard key frame set and a first user key frame set; Step S303, performing key frame matching on the first standard key frame set and the first user key frame set by using a dynamic time warping algorithm to obtain a second standard key frame set and a second user key frame set with the same number of key frame images; Step S304, calculating and obtaining an action completion score according to the second standard key frame set and the second user key frame set; Step S305, constructing a spatiotemporal network graph according to the second user key frame set; Step S306, inputting the spatiotemporal network graph into the scoring model and outputting the action fluency score; Step S307, performing weighted summation on the action completion score and the action fluency score to obtain a comprehensive score; Step S308, traverse and calculate the angle deviation values between all the skeletal key point lines in the key frame images of the second standard key frame set and the second user key frame set, and determine that the angle deviation value is greater than or equal to the angle deviation threshold, then mark the skeletal key point lines in the key frame image and return them to the user.