Gesture frame sequence generation method and device
By acquiring gesture frames from the video of the action to be processed and using a gesture classification model for automated classification and joint rotation information generation, the stability and consistency issues of gesture animation generation in existing technologies are solved, improving generation efficiency and smoothness, and making it suitable for the design of diverse digital cultural products.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUHAI KINGSOFT ONLINE GAME TECH CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, gesture animation generation based on motion capture is easily affected by the external environment, resulting in insufficient stability and consistency. Meanwhile, manual keyframe production is inefficient and difficult to adapt to diverse gesture requirements, leading to poor quality of generated character gesture animations.
By acquiring gesture frames from the video of the action to be processed, an automated classification model is used to construct an initial labeled gesture frame sequence, and a target gesture frame sequence is generated based on joint rotation information, thereby reducing background interference and improving the accuracy and fluency of classification information.
It achieves both improved generation efficiency and ensures the accuracy and smoothness of gesture movements, resulting in character gesture animations that better meet game quality requirements and are suitable for diverse digital cultural product designs.
Smart Images

Figure CN121904243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of digital cultural and creative activities and gesture frame sequence generation technology, and particularly to a method for generating gesture frame sequences. This application also relates to a gesture frame sequence generation apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the advancement of computer vision technology and the continuous progress of game development techniques, character animation is playing an increasingly important role in enhancing player immersion and interactive experience. Among these, character gesture animation, as a crucial form of character behavior expression, directly impacts the overall performance of the game through its realism and consistency.
[0003] Motion capture-based solutions automatically generate corresponding animation sequences by collecting human gesture information. However, the generated results are easily affected by external environmental conditions, resulting in insufficient animation stability and consistency. Manual keyframe creation, on the other hand, relies too heavily on human experience, is inefficient, and struggles to flexibly adapt to diverse gesture requirements. Therefore, a gesture frame sequence generation method is urgently needed to address these technical problems. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method for generating gesture frame sequences. This application also relates to a gesture frame sequence generation apparatus, a computing device, a computer-readable storage medium, and a computer program product, to solve the aforementioned problems existing in the prior art.
[0005] According to a first aspect of the embodiments of this application, a method for generating a gesture frame sequence is provided, comprising: Obtain at least one gesture frame from the video of the action to be processed, and determine the classification information corresponding to each gesture frame. At least one initial gesture frame is determined from each gesture frame to be processed based on the classification information. An initial labeled gesture frame sequence is constructed based on each initial gesture frame and the classification information corresponding to each initial gesture frame. The initial labeled gesture frame sequence is adjusted according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence; The target gesture frame sequence is generated based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
[0006] According to a second aspect of the embodiments of this application, a gesture frame sequence generation apparatus is provided, comprising: The classification module is configured to acquire at least one gesture frame to be processed in the action video to be processed, and determine the classification information corresponding to each gesture frame to be processed. The construction module is configured to determine at least one initial gesture frame from each gesture frame to be processed based on each classification information, and to construct an initial labeled gesture frame sequence based on each initial gesture frame and the classification information corresponding to each initial gesture frame. The adjustment module is configured to adjust the initial labeled gesture frame sequence according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence; The generation module is configured to generate a target gesture frame sequence based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
[0007] According to a third aspect of the embodiments of this application, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described gesture frame sequence generation method.
[0008] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described gesture frame sequence generation method.
[0009] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described gesture frame sequence generation method.
[0010] The gesture frame sequence generation method provided in this application involves: acquiring at least one gesture frame to be processed from a video of an action to be processed, and determining the classification information corresponding to each gesture frame; determining at least one initial gesture frame from each gesture frame to be processed based on the classification information; constructing an initial labeled gesture frame sequence based on each initial gesture frame and the classification information corresponding to each initial gesture frame; adjusting the initial labeled gesture frame sequence based on the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence; and generating a target gesture frame sequence based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
[0011] One embodiment of this application implements the generation of gesture frame sequences. It determines the gesture frames to be processed from the video of the action to be processed, reducing the impact of gesture-irrelevant regions on the classification information output and improving the accuracy of the classification information. Then, based on the classification information, at least one initial gesture frame is determined from each of the gesture frames to be processed. Initial labeled gesture frame sequences are constructed by retaining the initial gesture frames with high reliability of the classification results. The initial labeled gesture frame sequences are then adjusted according to the frame information and classification information corresponding to each initial gesture frame, so that the generated reference labeled gesture frame sequence can better reflect the change process of the gesture posture. A target gesture frame sequence is generated based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence. The joint rotation information reflects the gesture action process, improving the smoothness of the gesture frame sequence. The above-described method for generating gesture frame sequences can balance generation efficiency, action accuracy, and smoothness, making it better applicable to different project development and application scenarios. Attached Figure Description
[0012] Figure 1 This is a flowchart of a gesture frame sequence generation method provided in an embodiment of this application; Figure 2 This is a schematic diagram of various preset gesture types provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a method for generating gesture frame sequences in a game scene, provided in one embodiment of this application. Figure 4 This is a schematic diagram of the structure of a gesture frame sequence generation device provided in an embodiment of this application; Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this application. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0014] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application and in the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” used in one or more embodiments of this application refers to and includes any or all possible combinations of one or more associated listed items.
[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this application, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0017] First, the terms and concepts involved in one or more embodiments of this application will be explained.
[0018] Computer Vision (CV) is a branch of artificial intelligence that aims to enable computers to "see" and understand the world like humans. It processes and analyzes visual data such as images and videos to extract meaningful information, recognize patterns, make judgments, and take actions. Tasks it can perform include object recognition, tracking, and scene understanding. It utilizes technologies such as machine learning and deep learning to automatically extract information and make decisions from images, and its applications are wide-ranging.
[0019] Motion capture: Motion capture (Mocap) is a branch of computer vision that records and digitizes the movements of real people or objects. It captures motion trajectories and postures through sensors or cameras, converts real-world movements into digital data, and then applies it to computer-generated virtual models to achieve realistic human animation. It is widely used in scene construction in fields such as film, games, sports, medicine, and robotics to generate more human-like and realistic digital content or to conduct scientific analysis.
[0020] Region of Interest (ROI): A Region of Interest (ROI) is a specific area in an image or scene that is of interest or interest. For computer vision tasks, ROIs can be used to extract, manipulate, or analyze the features of that region.
[0021] Human body parametric model: A human body parametric model can be understood as a mathematical model that describes the structure and posture of the human body in the form of parameters. It can be used to model the shape and motion state of the human body while maintaining the topological relationships and structural constraints of the human skeleton. A human body parametric model typically includes body shape parameters to describe the body shape and posture parameters to describe the human posture. The posture parameters represent the motion state of each joint in the human body through joint rotation data, thereby achieving a unified representation of the human body's three-dimensional posture and movement.
[0022] Joint rotation sequences: Joint rotation sequences can be understood as a collection of joint rotation data arranged in chronological order, used to describe the posture changes of a human body or skeletal model over continuous time. Joint rotation sequences are represented parametrically, with each joint as a unit, relative to its parent joint. This is typically expressed using Euler angles, quaternions, or axis-angles. By combining joint rotation data from multiple time frames, the dynamic changes of human movement can be fully depicted.
[0023] Frame rate: Frame rate, also known as frame frequency, refers to the number of frames displayed per second. It is usually measured in FPS (Frames Per Second) or Hz (Hertz). It determines the smoothness and visual experience of a video or animation. The higher the frame rate, the smoother and more natural the picture; a low frame rate may result in stuttering or choppy playback.
[0024] In various fields of digital cultural and creative production and development, the realism and smoothness of character gesture animation directly affect players' perception of character movements and the overall interactive experience. With the continuous improvement of game graphics precision and interactive complexity, how to efficiently and stably generate high-quality gesture sequence frames has become one of the urgent technical problems to be solved in the production of digital cultural products and the design of digital cultural and creative products.
[0025] Existing gesture sequence frame generation schemes fall into two categories: one is a generation method based on real-time gesture recognition, which acquires human gesture data through acquisition devices such as cameras and maps the data into motion frames of a 3D character model; the other is a generation method based on manual keyframe creation, where animators manually set joint parameters corresponding to several key poses and then generate intermediate motion frames through interpolation calculations.
[0026] The aforementioned solutions all have certain limitations in practical applications. Specifically, real-time gesture recognition solutions typically rely on the external acquisition environment, and their generation quality is easily affected by factors such as lighting conditions, changes in shooting angle, and human posture occlusion. This leads to noise or instability in the acquired joint motion data, resulting in problems such as character posture distortion and skeletal misalignment, making it difficult to stably generate continuous gesture animations that meet game quality requirements. While manual keyframe creation can guarantee motion quality to a certain extent, the production process is highly dependent on the animator's experience, resulting in low production efficiency. Significant differences in production styles and joint settings among different animators lead to insufficient consistency in gesture movements, and it is also difficult to quickly adapt to the design and production needs of large-scale and diverse digital cultural products.
[0027] Furthermore, existing technologies mostly rely on discrete joint positions or static keyframes for animation generation, lacking a unified model of the internal structure of human movement. This makes it difficult to achieve continuous expression of gestures over time while ensuring the rationality of the skeletal structure. Therefore, there is an urgent need for a gesture frame sequence generation method that can improve generation efficiency while taking into account the accuracy, continuity, and natural fluency of gestures, thereby better meeting the application requirements for high-quality character animation in game scenarios.
[0028] This application provides a method for generating gesture frame sequences. This application also relates to a gesture frame sequence generating apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0029] The gesture frame sequence generation method and apparatus provided in this application can be used in digital culture production software and development. By acquiring the user-input video of the action to be processed, the target gesture frame sequence can be quickly generated, thereby realizing the production of digital culture products and providing users with more interesting digital culture content services.
[0030] Figure 1 The flowchart illustrates a gesture frame sequence generation method according to an embodiment of this application, which specifically includes the following steps: Step 102: Obtain at least one gesture frame to be processed from the action video to be processed, and determine the classification information corresponding to each gesture frame to be processed.
[0031] The action video to be processed can be understood as a real video of the user's body posture recorded by a camera device. It can be understood that the action video to be processed consists of at least one video frame to be processed. The gesture frame to be processed can be understood as a screenshot of the hand obtained frame by frame based on the video frame to be processed. The classification information can be understood as the output result of classifying and judging the gesture type of each gesture frame to be processed.
[0032] In practical applications, the acquired action video to be processed can be a real video of the user's body posture. However, the gesture posture itself is often small in size compared to the human torso. If the gesture posture is directly classified in the overall posture video, a lot of background information unrelated to the gesture posture will be introduced, thus affecting the accuracy of the classification information corresponding to each gesture frame to be processed.
[0033] Based on this, a preliminary analysis of the video of the action to be processed can be performed before classifying and judging the gesture type of each gesture frame to be processed. The candidate region containing the gesture action to be processed can be identified and obtained as the gesture frame to be processed, and then the classification information corresponding to each gesture frame to be processed can be identified.
[0034] In one specific embodiment of this application, the captured action video to be processed is parsed to obtain at least one gesture frame to be processed in the action video to be processed, and the classification information corresponding to each gesture frame to be processed is determined.
[0035] Taking the generation of gesture frames for game characters as an example, a real-time video containing the user's body posture is received as the action video to be processed, a total of 600 gesture frames to be processed are obtained, and the classification information corresponding to each gesture frame to be processed is determined.
[0036] By performing preliminary analysis on the video of the action to be processed, the candidate region containing the gesture frame to be processed is determined, thereby enabling more accurate classification in the subsequent classification stage. This method can effectively narrow the computational scope, reduce interference from background information, and achieve high-resolution data processing even in resource-constrained scenarios.
[0037] In one specific embodiment of this application, acquiring at least one gesture frame to be processed from a video of an action to be processed includes: Analyze the video of the action to be processed to obtain at least one initial video frame; Identify and extract the gesture region in each initial video frame to obtain at least one gesture frame to be processed.
[0038] The initial video frame can be understood as the constituent unit of the action video to be processed. The gesture region can be understood as the image region related to the gesture posture in each initial video frame, or as the region of interest in each initial video frame. The gesture frame to be processed is a video frame obtained by extracting the region of interest of each initial video frame.
[0039] The motion video to be processed is typically composed of multiple video frames. It can be sampled at a fixed or variable frame rate, with each frame corresponding to a time sampling point used to record the image information at that moment. By analyzing the motion video, continuous initial video frames can be obtained, which can be used to acquire the dynamic characteristics of the user's body posture, providing a foundation for subsequent digital cultural product production.
[0040] After acquiring each initial video frame, the hand gesture is used as the region of interest for extraction, which can enable more accurate classification in the subsequent gesture classification stage.
[0041] In one specific embodiment of this application, the action video to be processed is parsed to obtain at least one initial video frame, and the gesture region in each initial video frame is identified and extracted to obtain at least one gesture frame to be processed.
[0042] Taking a real-time video containing human gestures as an example, the frame rate of the video to be processed is 30fps and the resolution is 1920×1080. The video to be processed can be parsed by a computer vision processing library to obtain 600 initial video frames. The gesture region corresponding to the initial video frame is obtained frame by frame, and 600 gesture frames to be processed can be obtained accordingly.
[0043] By performing preliminary analysis on the video of the action to be processed, the candidate region containing the gesture frame to be processed is determined, thereby enabling more accurate classification in the subsequent classification stage. This video analysis and video frame extraction method, which ranges from coarse-grained broad range to fine-grained precise range, can effectively reduce the computational scope, reduce interference from background information, and achieve high-resolution data processing even in resource-constrained scenarios.
[0044] In a specific embodiment provided in this application, determining the classification information corresponding to each gesture frame to be processed includes: Each gesture frame to be processed is input into the gesture classification model to obtain the classification information corresponding to each gesture frame to be processed output by the gesture classification model.
[0045] A gesture classification model can be understood as an artificial intelligence model used to classify and determine the type of input gesture frames. It analyzes and calculates the input gesture frames, mapping the input data to gesture categories in a predefined set of gesture categories, thereby achieving type recognition or state determination of the gesture frames. Gesture classification models can be trained using supervised learning methods. By utilizing training data labeled with gesture categories, the model can determine the gesture category of input video frames or their feature data.
[0046] In classification scenarios involving computer vision, the amount of data in the gesture frames to be processed is usually large. Relying on manual methods to classify gestures is not only inefficient, but also difficult to maintain stable operation in large-scale cultural product design and development scenarios. Furthermore, manual classification is often influenced by subjective judgment and differences in experience; different people may give different classification results under the same input conditions, making it difficult to guarantee the consistency of classification results.
[0047] Therefore, automated classification can be achieved by employing a gesture classification model. This allows for batch analysis and category determination of input data without requiring manual intervention, thereby improving overall processing efficiency and meeting the system's requirements for real-time or large-scale processing. Automated classification allows for processing of input data based on unified classification rules and model parameters, ensuring stable classification results under identical conditions and improving the consistency and repeatability of the system's overall output.
[0048] In one specific embodiment of this application, each gesture frame to be processed is input into a gesture classification model to obtain the classification information corresponding to each gesture frame to be processed output by the gesture classification model.
[0049] Following the previous example, each gesture frame to be processed is input into the trained gesture classification model to obtain the classification information corresponding to each gesture frame.
[0050] By employing a gesture classification model, automatic analysis and category determination of video frames or video frame sequences are achieved, reducing manual intervention and improving overall processing efficiency. This approach is more suitable for processing large-scale video data. The automated classification method processes input data based on a unified model and decision rules, avoiding inconsistencies in classification results caused by differences in human experience. This helps ensure the stability of classification results under different scenarios and data conditions. The generated classification information can be directly used as input for subsequent gesture sequence frame generation and animation-driven processes, facilitating the construction of continuous and automated processing workflows and improving the overall system efficiency and practicality.
[0051] In one specific embodiment of this application, the gesture classification model is trained and generated through the following steps: Obtain an initial video sample set corresponding to at least one preset gesture type, wherein the initial video sample set includes at least one initial video sample corresponding to a video viewpoint; Determine at least one gesture image sample corresponding to each initial video sample, input each gesture image sample into the gesture classification model, and obtain the training classification information output by the gesture classification model; The model loss information is calculated based on the training classification information and the preset gesture type. The model parameters of the gesture classification model are adjusted according to the model loss information, and the gesture classification model is trained until the model training stops.
[0052] Preset gesture types can be understood as label information for training the gesture classification model. They can be used to indicate the gesture type to which a gesture image sample belongs and serve as the basis for the gesture classification model to learn the mapping relationship between input samples and target categories. The initial video sample set can be understood as a set consisting of at least one initial video sample corresponding to each preset gesture type. Each initial video sample is obtained based on a different video perspective. The video perspective can be understood as a common application perspective obtained using a shooting device. For game scenes, the video perspective can include front, side, and oblique perspectives, etc.
[0053] The initial video samples can be understood as original gesture video samples taken from different video perspectives based on a preset gesture type. Gesture image samples can be understood as gesture frame samples obtained by parsing the initial video samples. Training classification information can be understood as the classification results generated by the gesture classification model in the current training cycle. Training classification information can include the training classification results and training output confidence scores corresponding to each gesture image sample.
[0054] Model loss information can be understood as the loss value generated by the gesture classification model in the current training cycle. Model parameters can be understood as the parameters that can be adjusted based on the model loss information during the training process of the gesture classification model. The model training stopping condition can be understood as the preset judgment rule used to determine whether the training process should terminate during the model training process. It can be determined based on the number of training rounds, the magnitude of parameter updates, changes in loss values, or other indicators that can reflect the convergence state of the model. When the model training stopping condition is detected, the training process ends, and the gesture classification model retains its current parameter state as the training result.
[0055] During training, the model loss information can be used to determine the learning direction and optimization goal of the gesture classification model. By providing feedback through the prediction results output by the gesture classification model in the current state, the direction of parameter adjustment for the gesture classification model can be determined, and the strategy and parameters can be gradually adjusted and optimized in continuous training iterations.
[0056] The quality of initial video samples captured by a camera can be affected by various external factors, such as camera hardware quality, ambient lighting, shooting angle, and human posture occlusion. These factors can all influence the training results of a gesture classification model. If the model's accuracy is poor, it will affect the accuracy and reliability of classification information during application. Therefore, by capturing multi-view videos of preset gesture types, initial video samples from different perspectives can be obtained, enriching the number of samples in the training set for the gesture classification model. Furthermore, to ensure data diversity, the more gesture image samples obtained, the larger the amount of data used for model training, thereby improving data diversity and model accuracy.
[0057] In one specific embodiment of this application, at least one initial video sample set corresponding to a preset gesture type is obtained, wherein the initial video sample set includes at least one initial video sample corresponding to a video viewpoint, at least one gesture image sample corresponding to each initial video sample is determined, each gesture image sample is input into a gesture classification model, training classification information output by the gesture classification model is obtained, model loss information is calculated based on the training classification information and the preset gesture type, the model parameters of the gesture classification model are adjusted based on the model loss information, and the gesture classification model is trained until the model training stop condition is reached.
[0058] In one specific embodiment of this application, before training the gesture classification model, gesture image samples from each initial video sample can be identified and extracted as ROI regions. The ROI regions can also be preprocessed, such as normalization, cropping, and flipping enhancement, to improve the consistency of the input sample format and thus improve the model accuracy of the gesture classification model.
[0059] Taking the ResNet deep neural network model as an example for training as a gesture classification model, see [link to relevant documentation]. Figure 2 , Figure 2 A schematic diagram of various preset gesture types provided according to an embodiment of this application is shown. The preset gesture types include: three fingers raised, three fingers spread, antler gesture, orchid gesture, Buddha hand gesture, and flat palm gesture, totaling six types. Initial video sample sets corresponding to the six preset gesture types are obtained, each initial video sample set including initial video samples corresponding to front, side, and oblique perspectives, respectively. It is understood that the above-mentioned preset gesture types and video perspectives are illustrative examples of gesture types and video perspectives; in actual applications, gesture types and video perspectives can be set according to the project content and actual scene characteristics.
[0060] Next, each initial video sample was analyzed. 1500 initial video frame samples were obtained from different video perspectives of the 6 preset gesture types. A total of 27,000 initial video frame samples were obtained. The initial video frame samples were preprocessed to crop out hand ROI regions with uniform format and normalize them. The pixel values were mapped to the range [0, 1] and horizontal flip enhancement was applied. Finally, 27,000 gesture image samples were obtained.
[0061] The dataset can be divided into 80% training set and 20% test set. The model training stops when the classification accuracy is greater than or equal to 95%. The model loss is calculated based on the training classification information output by the ResNet network and the preset gesture type, and the model parameters are adjusted. The model training stops when the classification accuracy is greater than or equal to 95%.
[0062] By collecting initial video samples from different perspectives, complementary information can be obtained from various angles, helping the model learn more perspectives and more stable discriminative features, thereby improving classification performance and effectively increasing the accuracy of the gesture classification model. Since the initial video samples are collected from different perspectives, the subsequently generated frame sequences can be adapted to different application scenarios and perspectives, making them more versatile.
[0063] Step 104: Determine at least one initial gesture frame from each gesture frame to be processed based on the classification information, and construct an initial labeled gesture frame sequence based on each initial gesture frame and the classification information corresponding to each initial gesture frame.
[0064] The initial gesture frame can be understood as the gesture frame to be processed obtained by filtering according to classification information. The initial gesture frame can include gesture image information and time series information. Among them, gesture image information can represent the gesture posture and the preset gesture type to which the gesture posture belongs, and time series information can represent the playback order of each initial gesture frame. It can provide a sorting basis for the construction of the subsequent initial labeled gesture frame sequence, thereby forming a continuous dynamic visual effect and a smooth playback experience. The initial labeled gesture frame sequence can be understood as the gesture frame sequence composed of the remaining initial gesture frames after filtering each initial gesture frame according to time series information and classification information and removing gesture frames that belong to noise.
[0065] A gesture frame to be processed can be understood as a unit of a video of an action to be processed. Such videos are typically composed of a large number of gesture frames. Directly constructing a sequence of labeled gesture frames based on these frames and their classification information is complex and time-consuming. To improve the efficiency of constructing the gesture frame sequence, we can select gesture frames with significant gesture features based on classification information to construct an initial sequence of labeled gesture frames.
[0066] In a specific embodiment provided in this application, each gesture frame to be processed is filtered according to the classification information of each gesture frame to be processed, at least one initial gesture frame is determined, and an initial labeled gesture frame sequence is constructed according to each initial gesture frame and the classification information corresponding to each initial gesture frame.
[0067] Taking 5000 gesture frames to be processed as an example, we obtain the 5000 gesture frames to be processed and the classification information corresponding to each gesture frame. Based on the classification information of each gesture frame to be processed, we filter the gesture frames to be processed and determine 25 initial gesture frames. Based on each initial gesture frame and the classification information corresponding to each initial gesture frame, we construct an initial labeled gesture frame sequence.
[0068] The method of first filtering each gesture frame to be processed based on the classification information of the gesture frames to be processed, and then constructing the initial labeled gesture frame sequence, can improve the construction accuracy and efficiency of the initial labeled gesture frame sequence.
[0069] In one specific embodiment provided in this application, the classification information includes classification results and classification confidence levels; Based on the classification information, at least one initial gesture frame is determined from each gesture frame to be processed. Based on each initial gesture frame and the classification information corresponding to each initial gesture frame, an initial labeled gesture frame sequence is constructed, including: Each gesture frame to be processed with a classification confidence score greater than or equal to the classification confidence threshold is used as the initial gesture frame; Based on each initial gesture frame and the corresponding classification result, construct the first labeled gesture frame sequence; The first labeled gesture frame sequence is updated based on the frame interval between adjacent initial gesture frames and the denoising frame interval threshold in the first labeled gesture frame sequence, and the initial labeled gesture frame sequence is constructed.
[0070] The classification result can be understood as the gesture type of each gesture frame to be processed. It can be understood that the gesture type of each gesture frame to be processed corresponds to the preset gesture type. The gesture type can be preset according to the application scenario. The classification confidence can be understood as the credibility of the output classification result. The classification confidence threshold can be understood as the preset classification confidence filtering criterion.
[0071] The first labeled gesture frame sequence can be understood as a sequence of gesture frames obtained by filtering each initial gesture frame according to the classification confidence level and classification confidence threshold, removing gesture frames with a classification confidence level lower than the classification confidence threshold, and labeling the initial gesture frames with a classification confidence level greater than or equal to the classification confidence threshold and their corresponding classification results. Accordingly, the first labeled gesture frame sequence includes the motion change information of the gesture and the gesture type. The gesture type is determined according to the classification result in the classification information. Accordingly, adjacent initial gesture frames can be understood as initial gesture frames that are adjacent in the first labeled gesture frame sequence according to the time series information.
[0072] The frame interval can be understood as the number of frames between two adjacent initial gesture frames. The denoising frame interval can be understood as a pre-set frame interval threshold, which can be determined based on the action time required to switch from one preset gesture type to another. If the frame interval is less than the denoising frame interval, it can be determined that there is noise in the adjacent initial gesture frames, and the noisy gesture frames need to be updated and removed before constructing the initial labeled gesture frame sequence. If the frame interval is greater than or equal to the denoising frame interval, it can be determined that there is no noise in the adjacent initial gesture frames, and the adjacent initial gesture frames are retained.
[0073] After classifying the input gesture frames, the gesture classification model outputs the same number of results as the number of gesture frames, and this number is usually quite large, inevitably containing noise or misclassifications. If the output of the gesture classification model is not filtered and directly used to construct subsequent gesture frame sequences, it will not only increase the data scale of subsequent processing and reduce processing efficiency, but may also affect the accuracy of the final processing results due to the introduction of noise data.
[0074] Therefore, the initial gesture frames output by the gesture classification model can be filtered from different dimensions. These dimensions can include filtering dimensions based on the model's output results and filtering dimensions based on the attributes of the initial gesture frames themselves. The gesture classification model can output confidence scores along with the gesture classification results. The classification confidence score and classification confidence threshold can filter out initial gesture frames with classification confidence scores below the classification confidence threshold, retaining initial gesture frames with classification confidence scores greater than or equal to the classification confidence threshold as valid output results, thereby constructing the first labeled gesture frame sequence.
[0075] Based on the construction of the first labeled gesture frame sequence, filtering can be performed according to the time series information between each initial gesture frame in the first labeled gesture frame sequence. If the interval is less than the denoised frame interval, it can be determined that there is noise in the adjacent initial gesture frames. This is because within the denoised frame interval threshold, the user cannot complete the change between two preset gesture types, so it is necessary to update and remove the noisy gesture frames before constructing the initial labeled gesture frame sequence. If the interval is greater than or equal to the denoised frame interval, it can be determined that there is no noise in the adjacent initial gesture frames, and the adjacent initial gesture frames are retained.
[0076] In a specific embodiment provided in this application, each gesture frame to be processed with a classification confidence level greater than or equal to the classification confidence threshold is taken as an initial gesture frame, and a first labeled gesture frame sequence is constructed based on each initial gesture frame and the classification result corresponding to each initial gesture frame. Obtain the frame interval and denoising frame interval threshold between adjacent initial gesture frames in the first labeled gesture frame sequence. If the frame interval between adjacent initial gesture frames is less than the denoising frame interval threshold, and it is assumed that the initial gesture frame with the earlier time sequence in the time series information of the adjacent initial gesture frames is retained, then the initial gesture frame with the later time sequence in the time series information of the adjacent initial gesture frames is a noise gesture frame and can be deleted. Correspondingly, if it is assumed that the initial gesture frame with the later time sequence in the time series information of the adjacent initial gesture frames is retained, then the initial gesture frame with the earlier time sequence in the time series information of the adjacent initial gesture frames is a noise gesture frame and can be deleted.
[0077] If the frame interval between adjacent initial gesture frames is greater than or equal to the denoising frame interval threshold, it can be determined that there is no noise in the adjacent initial gesture frames, and the adjacent initial gesture frames are retained.
[0078] In the specific embodiments provided in this application, the method for determining noisy gesture frames in adjacent initial gesture frames is not limited. As long as the frame interval between adjacent initial gesture frames in the first labeled gesture frame sequence is less than the noise reduction frame interval threshold, the noisy gesture frames can be deleted.
[0079] Taking a classification confidence threshold of 85% as an example, gesture frames with a classification confidence of less than 85% are excluded, and initial gesture frames with a classification confidence of greater than or equal to 85% are retained. Based on each initial gesture frame and the classification result corresponding to each initial gesture frame, the first labeled gesture frame sequence is constructed.
[0080] Correspondingly, if the classification result of another gesture frame to be processed is a natural pose with a classification confidence of 60%, which is less than the classification confidence threshold of 85%, then the gesture frame to be processed is excluded. The above method is used to filter each gesture frame to be processed one by one to obtain the first labeled gesture frame sequence.
[0081] If the first labeled gesture frame sequence is “Natural-clenched fist-orchid-natural”, and the noise reduction frame interval threshold is 6 frames, the frame interval between adjacent gesture frames in the first labeled gesture frame sequence is obtained as follows: “Natural-clenched fist” is 2 frames; “clenched fist-orchid” is 12 frames; “orchid-natural” is 12 frames.
[0082] The frame interval for "Natural-Fist" is less than the denoising frame interval threshold. Assuming we retain the earlier "Natural" frame, then the "Fist" frame in the adjacent initial gesture frames is a noisy gesture frame and can be deleted. The frame intervals for "Fist-Orchid" and "Orchid-Natural" are greater than the denoising frame interval threshold, so they do not need to be adjusted and can be retained. The initial labeled gesture frame sequence is "Natural-Orchid-Natural," where the frame interval for "Natural-Orchid" is adjusted to 14 frames, and the frame interval for "Orchid-Natural" is 12 frames.
[0083] The initial gesture frames output by the gesture classification model are filtered using different dimensions. Filtering is based on classification confidence thresholds and the classification confidence of each gesture frame to be processed, retaining valid initial gesture frames for subsequent frame sequence construction. Further filtering is performed on the inherent attributes of the initial gesture frames to construct the initial labeled gesture frame sequence. These filtering steps remove invalid or redundant gesture frames, improving the efficiency and accuracy of constructing the initial labeled gesture frame sequence.
[0084] Step 106: Adjust the initial labeled gesture frame sequence according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence.
[0085] Frame information can be understood as the attribute information of each initial gesture frame, which can include the gesture image information and time sequence information of each initial gesture frame. Among them, the gesture image information can be used to represent the gesture posture, and the time sequence information can represent the playback order of each initial gesture frame. The reference annotation gesture frame sequence can be understood as the gesture frame sequence obtained by interpolating the initial annotation gesture frame sequence according to the sorting rules. The gesture frame subsequence can be understood as the subsequence that transitions from the initial type gesture to the functional type gesture and then back to the initial type gesture. Accordingly, the gesture frame subsequence can include a start node, an intermediate node, and an end node. The gesture type corresponding to each node can be preset according to the preset gesture type. It can be set that the start node and end node are the initial type gestures, and the intermediate node is the functional type gesture. The gesture frame subsequence can represent the action process of how to reasonably transition from the natural state of the initial type to the specific posture of the functional type.
[0086] Gesture frame subsequences are used to reflect the process of a user transitioning from a natural state to a specific posture when performing a specific gesture. Therefore, between specific postures of a function type, a natural posture of the initial type is needed for connection and transition. However, in the initial labeled gesture frame sequence obtained by filtering based on classification confidence, there may be cases where the natural state is missing between specific postures of the initial labeled gesture sequence. This missing state may be caused by a variety of factors. Therefore, the initial labeled gesture frame sequence needs to be adjusted to obtain a reference labeled gesture frame sequence, which may include each gesture frame subsequence.
[0087] In a specific embodiment provided in this application, the initial labeled gesture frame sequence is adjusted according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence.
[0088] Taking the classification information of each initial gesture frame in the initial labeled gesture frame sequence as natural, fist, orchid, and natural as an example, the initial labeled gesture frame sequence is adjusted according to the frame information of each initial gesture frame to obtain a reference labeled gesture frame sequence, which includes at least one gesture frame subsequence.
[0089] Adjusting the initial labeled gesture frame sequence according to the sorting rules, the resulting reference labeled gesture frame sequence can better reflect the user's movement process from a natural state to a specific posture when making a specific gesture, which helps to improve the generation quality of the target gesture frame sequence.
[0090] In a specific embodiment provided in this application, the classification information includes classification results and classification confidence, wherein the classification results include initial type and at least one functional type; The initial labeled gesture frame sequence is adjusted based on the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, including: At least one adjacent initial gesture frame is determined based on the frame information corresponding to each initial gesture frame. Determine the target adjacent initial gesture frame and the target frame interval of the target adjacent initial gesture frame, wherein the target adjacent initial gesture frame is a function type; If the target frame interval is greater than or equal to the frame interval threshold, a reference annotation gesture frame is determined based on the target's adjacent initial gesture frame and the initial gesture frame of the starting type. If the target frame interval is less than the frame interval threshold, a reference labeled gesture frame is determined based on the classification confidence of the target's adjacent initial gesture frames; The reference annotation gesture frame sequence is constructed based on each reference annotation gesture frame.
[0091] The starting type can be understood as a natural gesture posture, the function type can be understood as the gesture posture corresponding to the preset function, the adjacent initial gesture frames can be understood as initial gesture frames that are determined to be adjacent based on the time sequence relationship, the target adjacent initial gesture frames can be understood as initial gesture frames whose classification result is the function type and whose time sequence relationship is adjacent, or they can be understood as adjacent initial gesture frames that need to be optimized and adjusted, the target frame interval can be understood as the frame interval between the target adjacent initial gesture frames, the frame interval threshold can be understood as the preset interval threshold, and the reference annotation gesture frame can be understood as the gesture frame used to construct the reference annotation gesture frame sequence. The way to generate the reference annotation gesture frame also varies depending on the relationship between the target frame interval and the frame interval threshold.
[0092] Since the initial gesture frame sequence is determined based on whether it meets the classification confidence threshold, without filtering the classification results of the initial gesture frame sequence to conform to the generated frame sequence of the preset arrangement order, it is possible to decide whether to perform secondary optimization and adjustment based on the classification results, and then decide the adjustment method based on the relationship between the target frame interval and the frame interval threshold, so as to ensure that the generated reference annotation gesture frame sequence includes the frame sequence of gestures of various function types with natural posture as the transition posture.
[0093] In a specific embodiment provided in this application, at least one adjacent initial gesture frame is determined based on the frame information corresponding to each initial gesture frame; a target adjacent initial gesture frame and a target frame interval of the target adjacent initial gesture frame are determined, wherein the target adjacent initial gesture frame is a function type; if the target frame interval is greater than or equal to a frame interval threshold, a reference annotation gesture frame is determined based on the target adjacent initial gesture frame and the initial gesture frame of the starting type; if the target frame interval is less than the frame interval threshold, a reference annotation gesture frame is determined based on the classification confidence of the target adjacent initial gesture frame; and a reference annotation gesture frame sequence is constructed based on each reference annotation gesture frame.
[0094] Using the previous example, if the classification information of each initial gesture frame in the initial labeled gesture frame sequence is "natural, fist, orchid, natural", then it can be determined that there are 3 adjacent initial gesture frames, including "natural-fist"; "fist-orchid"; "orchid-natural"; among which the target's adjacent initial gesture frame is "fist-orchid". If the frame interval between adjacent initial gesture frames of the target is 3 ps (picosecond) and the frame interval threshold is 2 ps, then the reference annotation gesture frames are determined based on the adjacent initial gesture frames of the target and the initial gesture frames of the starting type, and a reference annotation gesture frame sequence is constructed. If the frame interval between adjacent initial gesture frames of the target is 1ps and the frame interval threshold is 2ps, then the reference annotation gesture frames are determined based on the classification confidence of adjacent initial gesture frames of the target, and a reference annotation gesture frame sequence is constructed. The adjustment method for the initial labeled gesture frame sequence is determined based on the frame interval and frame interval threshold. The obtained reference labeled gesture frame sequence can better reflect the action process from one preset gesture type to another preset gesture type, which helps to improve the generation quality of subsequent target gesture frame sequences.
[0095] In a specific embodiment provided in this application, determining a reference annotation gesture frame based on the target adjacent initial gesture frame and the initial gesture frame includes: Insert an initial gesture frame of the starting type into the adjacent initial gesture frames of the target to obtain a reference annotation gesture frame.
[0096] To better reflect the movement process from the initial type of natural hand gesture to the specific hand gesture of each function type, the video of the action to be processed starts with the natural posture as the starting point and then transitions to the specific posture. Therefore, complete action transition videos were recorded for hand gestures of different function types.
[0097] If the target frame interval is greater than or equal to the frame interval threshold, and the adjacent initial gesture frames of the target are all of the functional type, the possible factor is the loss of the intermediate natural pose. In this case, an initial gesture frame of the starting type can be inserted into the adjacent initial gesture frames of the target to generate a reference annotation gesture frame.
[0098] In a specific embodiment provided in this application, if the target's adjacent initial gesture frames are all of function type and the frame interval is greater than a threshold, then an initial gesture frame of the starting type is inserted into the target's adjacent initial gesture frames to generate a reference annotation gesture frame.
[0099] Taking the initial gesture frames adjacent to the target as a fist and an orchid as an example, the target frame interval is 5ps and the frame interval threshold is 3ps. Then, an initial gesture frame of the starting type is inserted between the fist and the orchid. The reference labeled gesture frame sequence is "natural-fist-natural-orchid-natural".
[0100] When the target frame interval is greater than or equal to the frame interval threshold, the missing initial gesture frames of the starting type are automatically inserted and supplemented. The generated reference annotation gesture frames can better reflect the change process of the action and improve the accuracy of the subsequently constructed frame sequence.
[0101] In one specific embodiment of this application, determining a reference annotation gesture frame based on the classification confidence of the target's adjacent initial gesture frames includes: The classification confidence scores of each adjacent initial gesture frame are compared to obtain the comparison results; The initial gesture frame to be retained is determined based on the comparison results, and the reference annotation gesture frame is generated.
[0102] The comparison result can be understood as a comparison of the classification confidence levels of adjacent initial gesture frames. The comparison result can be used to identify and retain the initial gesture frame with higher confidence among adjacent initial gesture frames.
[0103] Since the motion video to be processed can include the motion process from a natural pose of the initial type to specific poses corresponding to various function types, if the classification results of adjacent initial gesture frames are still identified as function types even when the frame interval is less than the frame interval threshold, it may be due to errors in the classification results. This is because the time required to complete the transition from the initial pose to a specific pose and back to the initial pose within a short time interval is relatively long. If the frame interval is less than the frame interval threshold, it is impossible to achieve the aforementioned changes in specific poses for multiple function types, and it is more likely that the classification results in the classification information are incorrect. Therefore, the classification confidence of each adjacent initial gesture frame can be compared, and the initial gesture frame with the higher classification confidence can be retained to generate reference annotation gesture frames.
[0104] In a specific embodiment provided in this application, the classification confidence of each adjacent initial gesture frame is compared to obtain the comparison result. Based on the comparison result, the initial gesture frame to be retained is determined, and a reference labeled gesture frame is generated.
[0105] Taking the target's adjacent initial gesture frame as "clenched fist - orchid" as an example, the target frame interval is 5ps, the frame interval threshold is 10ps, the classification confidence of clenched fist is 73%, and the classification confidence of orchid is 87%. Then, the initial gesture frame of orchid is retained, and the initial gesture frame of clenched fist is deleted. The reference labeled gesture sequence is "natural - orchid - natural".
[0106] Determining the initial gesture frames to retain based on classification confidence ensures that the generated reference gesture frames have high classification confidence, thereby ensuring the reliability of the subsequently generated reference gesture frame sequence.
[0107] Step 108: Generate the target gesture frame sequence based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
[0108] Joint rotation information can be understood as the motion process information of hand joints when changing different hand gesture postures. Joint rotation information can generate interpolation data corresponding to each gesture frame subsequence. The target gesture frame sequence can be understood as the complete frame sequence obtained by interpolating each gesture subsequence based on joint rotation information.
[0109] In sequence-based gesture generation, simply acquiring reference marker gesture frames from each gesture frame subsequence lacks a depiction of the intermediate changes in the movement, making it difficult to reflect the continuous evolution of the gesture over time. This can easily lead to discontinuities and unnaturalness in the generated target gesture frame sequence regarding posture changes, motion trajectories, and rhythm, thus affecting the accuracy and realism of the gesture expression. Therefore, interpolation can be performed based on the joint rotation information corresponding to each gesture frame subsequence. The generated target gesture frame sequence includes not only the nodes of each gesture frame's posture but also information about the process of gesture changes.
[0110] In a specific embodiment provided in this application, joint rotation information corresponding to each gesture frame sub-sequence is obtained, and a target gesture frame sequence is generated based on each gesture frame sub-sequence and the joint rotation information corresponding to each gesture frame sub-sequence.
[0111] Taking the gesture frame subsequence as "Natural-Orchid" as an example, the joint rotation information from the natural posture to the orchid posture is determined, and the target gesture frame sequence is generated based on the gesture frame subsequence and the joint rotation information corresponding to the gesture frame subsequence.
[0112] Interpolation is performed based on the joint rotation information corresponding to each gesture frame subsequence. The generated target gesture frame sequence includes not only the nodes of each gesture frame's posture but also information about the gesture change process. This method is simple and quick. Furthermore, because the joint rotation information maintains the parent-child relationship constraint between hand joints, the joint rotation during the intermediate transition process conforms to the laws of gesture movement. Consequently, the generated target gesture frame sequence is complete and smooth.
[0113] In a specific embodiment provided in this application, a target gesture frame sequence is generated based on each gesture frame sub-sequence and the joint rotation information corresponding to each gesture frame sub-sequence, including: Determine a reference gesture frame subsequence, wherein the reference gesture frame subsequence is any one of the gesture frame subsequences; The target gesture frame is determined based on the classification information of the reference gesture frame subsequence, and the target joint rotation information is determined in the preset joint rotation information database. Generate a joint rotation frame sequence based on the reference gesture frame subsequence and the target joint rotation information; The target gesture frame sequence is generated based on the target gesture frame and the joint rotation frame sequence.
[0114] A reference gesture frame subsequence can be understood as any subsequence selected from various gesture frame subsequences. It can be used to explain the specific implementation method for generating the target gesture frame sequence. The target gesture frame can be used to characterize the gesture type corresponding to the reference gesture frame subsequence. The preset joint rotation information library can be understood as a pre-established data set used to store the action process of different gesture posture changes. The target joint rotation information can be understood as obtained from the preset joint rotation information library based on the classification information of the reference gesture frame subsequence. The joint rotation frame sequence, which matches the corresponding gesture type, can be understood as an interpolated frame sequence used to characterize the continuous transition process between gesture frames in the target gesture frame sequence.
[0115] The preset joint rotation information library stores a set of data on the motion process from a natural posture to a specific posture. This can be understood as including joint rotation information for different motion processes. However, since each joint rotation information is generated based on a preset reference hand model, and different users' hand structures differ in the number of joints, bone length ratios, and initial postures, directly applying the target joint rotation information from the preset library to generate the user's target gesture frame sequence can easily lead to motion distortion or posture mismatch. Therefore, before applying the target joint rotation information to the target hand model, it is necessary to adapt the joint rotation information according to the user's hand structure to ensure that the generated joint rotation frame sequence matches the user's hand model.
[0116] In a specific embodiment provided in this application, a reference gesture frame subsequence is determined; a target gesture frame is determined based on the classification information of the reference gesture frame subsequence; and target joint rotation information is determined in a preset joint rotation information database. A joint rotation frame sequence is generated based on the reference gesture frame subsequence and the target joint rotation information. The target gesture frame sequence is then generated based on the target gesture frame and the joint rotation frame sequence.
[0117] Taking "Nature-Orchid-Nature" as a reference gesture frame subsequence as an example, if the classification results in the classification information include Nature and Orchid, then the target gesture frames are determined to be Nature and Orchid. The target joint rotation information from Nature to Orchid and from Orchid to Nature is determined in the preset joint rotation information library. A joint rotation frame sequence that matches the user's hand model is generated. Then, the target gesture frame sequence is constructed based on the target gesture frames Nature and Orchid, as well as the joint rotation frame sequence from Nature to Orchid.
[0118] By selecting target joint rotation information that matches the target gesture frame from a pre-defined joint rotation information database and adapting it to the user's hand skeletal characteristics, the generated joint rotation frame sequence maintains consistency in parent-child constraints and movement style between joints, improving the stability and consistency of the motion generation results. Simultaneously, the motion generation process is transformed into a joint rotation sequence retrieval and adaptation process. For the parts requiring interpolation in each reference gesture frame subsequence, reasonable intermediate state rotation data can be easily and conveniently generated using quaternion interpolation. This method reduces computational complexity, improves system efficiency, and enhances the reusability of motion sequences and the system's scalability.
[0119] In a specific embodiment provided in this application, a joint rotation frame sequence is generated based on the reference gesture frame sub-sequence and the target joint rotation information, including: The number of supplementary frames is determined based on the number of frames in the reference gesture frame subsequence and the number of frames in the target gesture frame; A joint rotation frame sequence is generated based on the number of supplementary frames and the target joint rotation information.
[0120] The number of frames in the reference gesture frame subsequence can be understood as the number of frames required to complete the corresponding action change in the action video to be processed. The number of frames in the target gesture frame can be understood as the number of poses used to represent the target gesture, which can be determined based on the classification results. The number of supplementary frames can be used to represent the number of intermediate transition frames generated between adjacent key gesture frames. The supplementary frame number can be understood as the number of gesture types determined by the classification results, and can be used to represent the number of intermediate transition frames generated between adjacent key gesture frames.
[0121] By reasonably determining the number of supplementary frames, it can be ensured that the generated joint rotation frame sequence has visual continuity and naturalness.
[0122] Taking the reference gesture frame subsequence "Nature-Orchid-Nature" as an example, if the reference gesture frame subsequence includes a total of 60 frames and the corresponding target gesture frame sequence includes 3 target gesture frames, then the number of supplementary frames is determined to be 57. The resulting joint rotation frame sequence includes 3 target gesture frames and 57 supplementary frames inserted between the target gesture frames.
[0123] By generating a joint rotation frame sequence based on the number of supplementary frames, the generation efficiency of the joint rotation frame sequence can be improved while ensuring the smooth and natural movement of the target gesture frame sequence, thereby meeting the real-time interaction needs of various digital cultural and creative production processes.
[0124] In a specific embodiment provided in this application, the preset joint rotation information database is constructed through the following steps: Determine the starting gesture and at least one functional gesture; Obtain motion sampling videos corresponding to each functional gesture, wherein the motion sampling videos include the starting gesture and the target functional gesture, and the target functional gesture is any one of the functional gestures; Each motion sampling video is input into a rotation parameterization model, which outputs preset joint rotation information corresponding to each motion sampling video.
[0125] The starting gesture can be understood as a pre-set baseline gesture posture, such as a natural gesture posture. The functional gesture can be understood as the gesture posture corresponding to each specific function. The motion sampling video can be understood as a motion video describing the transition of the gesture from the starting gesture to the target functional gesture. The shooting angle of the motion sampling video can be consistent with the video angle. The target functional gesture is any one of the functional gestures, used to illustrate the construction of the motion sampling video.
[0126] The rotational parametric model can be understood as parsing the rotational data of each joint of the hand from motion sampling video, and parametrically processing the rotational data based on the parent-child relationship of the joints to obtain a parametric model containing the rotational information of the hand joints. Preset joint rotational information can be used to characterize the rotational features of the hand joints in the motion sampling video over time, and can serve as input for three-dimensional joint data in the subsequent motion generation process.
[0127] The pre-defined key rotation information library provides a unified standard for joint rotation data, reducing reliance on the creator's personal experience and improving the efficiency and consistency of gesture creation.
[0128] In a specific embodiment provided in this application, a starting gesture and at least one functional gesture are determined, and motion sampling videos corresponding to each functional gesture are obtained. The motion sampling videos include the starting gesture and the target functional gesture, where the target functional gesture is any one of the functional gestures. Each motion sampling video is input into a rotational parameterization model, and the rotational parameterization model outputs preset joint rotation information corresponding to each motion sampling video.
[0129] Taking four functional gestures—clenched fist, orchid, supporting hand, and hanging hand—as examples, 10 motion sampling videos were recorded for each of the four functional gestures, transitioning from natural to functional gesture and back to natural. Each video segment was 2 seconds long and contained 60 frames, for a total of 40 motion sampling videos. The motion sampling videos were then imported into a rotational parameterization model, and the quaternion data of the 21 joints of the hand in each 60 frames of the motion sampling video were parsed to form 40 sets of preset joint rotation information. Each set of preset joint rotation information contained 60 frames of data, and each frame contained 21 quaternions.
[0130] In generating preset joint rotation information, quaternion spherical linear interpolation is used in conjunction with skeletal constraints in the rotational parametric model to ensure that the joint rotation process conforms to the physiological movement patterns of the human body, avoiding stiff movements. Furthermore, this construction method has strong scalability; subsequent expansion of the preset joint rotation information library only requires adding new functional gestures and their corresponding motion sampling videos, eliminating the need to redesign the algorithm and reducing project development costs.
[0131] In one specific embodiment of this application, after constructing the initial labeled gesture frame sequence, the method further includes: A target gesture frame sequence is generated based on the initial labeled gesture frame sequence and the joint rotation information corresponding to the initial labeled gesture frame sequence.
[0132] Since the initial labeled gesture frame sequence has been filtered through different dimensions to remove noise, the target gesture frame sequence can be directly generated based on the initial labeled gesture frame sequence and its corresponding joint rotation information. This is because the joint rotation information can generate a corresponding quaternion sequence, which can be used for interpolation calculation.
[0133] In one specific embodiment of this application, a target gesture frame sequence is generated based on the initial labeled gesture frame sequence and the joint rotation information corresponding to the initial labeled gesture frame sequence.
[0134] Taking the initial labeled gesture frame sequence as "clenched fist - orchid" as an example, the joint rotation information of the corresponding postures of the clenched fist and orchid gestures can be obtained respectively, and the joint rotation information can be represented in the form of quaternions. Based on the initial labeled gesture frame sequence of "clenched fist - orchid" and its corresponding quaternion joint rotation information, the joint rotation information for transition is generated through interpolation, thereby generating the target gesture frame sequence.
[0135] In different application scenarios, target gesture frame sequences can be generated in different ways according to actual needs. As one implementation method provided in this application, when generating a target gesture frame sequence based on an initial labeled gesture frame sequence and its corresponding joint rotation information, quaternion interpolation can be directly performed on the joint rotation information to obtain the target gesture frame sequence. This generation method is simple to implement, has high computational efficiency, and can be flexibly configured according to actual needs, making it suitable for application scenarios with high real-time requirements.
[0136] One embodiment of this application implements the generation of gesture frame sequences. It determines the gesture frames to be processed from the video of the action to be processed, reducing the impact of gesture-irrelevant regions on the classification information output and improving the accuracy of the classification information. Then, based on the classification information, at least one initial gesture frame is determined from each of the gesture frames to be processed. Initial labeled gesture frame sequences are constructed by retaining the initial gesture frames with high reliability of the classification results. The initial labeled gesture frame sequences are then adjusted according to the frame information and classification information corresponding to each initial gesture frame, so that the generated reference labeled gesture frame sequence can better reflect the change process of the gesture posture. A target gesture frame sequence is generated based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence. The joint rotation information reflects the gesture action process, improving the smoothness of the gesture frame sequence. The above-described method for generating gesture frame sequences can balance generation efficiency, action accuracy, and smoothness, making it better applicable to different project development and application scenarios.
[0137] The following is in conjunction with the appendix Figure 3 Taking the application of the gesture frame sequence generation method provided in this application in a game scene as an example, the gesture frame sequence generation method will be further explained. Figure 3 This paper illustrates a flowchart of a gesture frame sequence generation method for a game scene according to an embodiment of this application, which specifically includes the following steps: Step 302: Determine the preset gesture type, which includes the starting type and the function type. The starting type includes the natural posture, and the function type includes the clenched fist, orchid finger, supporting hand, and hanging hand.
[0138] Step 304: Obtain the initial video samples corresponding to the four video perspectives for each preset gesture type and the gesture image samples corresponding to each initial video sample. Input each gesture image sample into the gesture classification model to train the gesture classification model.
[0139] Step 306: Determine the starting gesture, which includes natural posture gestures, and the functional gestures, which include clenched fist, orchid finger, supporting hand, and hanging hand. Obtain the motion sampling video corresponding to each functional gesture.
[0140] Step 308: Input the sampled videos of each action into the rotation parameterization model, and the rotation parameterization model outputs the preset joint rotation information corresponding to each sampled video of the action.
[0141] Step 310: Obtain at least one gesture frame to be processed from each initial video frame of the action video to be processed, input each gesture frame to be processed into the gesture classification model, and obtain the classification result and classification confidence of each gesture frame to be processed output by the gesture classification model.
[0142] Step 312: Take each gesture frame to be processed with a classification confidence score greater than or equal to the classification confidence threshold as the initial gesture frame. Construct a first labeled gesture frame sequence based on each initial gesture frame and the classification result corresponding to each initial gesture frame. Update the first labeled gesture frame sequence based on the frame interval between adjacent initial gesture frames and the denoising frame interval threshold in the first labeled gesture frame sequence to construct the initial labeled gesture frame sequence.
[0143] Step 314: Determine at least one adjacent initial gesture frame based on the frame information corresponding to each initial gesture frame, determine the target adjacent initial gesture frame and the target frame interval of the target adjacent initial gesture frame, wherein the target adjacent initial gesture frame is a function type.
[0144] Step 316: If the target frame interval is greater than or equal to the frame interval threshold, insert an initial gesture frame of the starting type into the target adjacent initial gesture frames to obtain a reference labeled gesture frame. If the target frame interval is less than the frame interval threshold, compare the classification confidence of each adjacent initial gesture frame, determine the initial gesture frame to be retained based on the comparison result, and generate the reference labeled gesture frame.
[0145] Step 318: Construct a reference annotation gesture frame sequence based on each reference annotation gesture frame, wherein the reference annotation gesture frame sequence includes at least one gesture frame subsequence.
[0146] Step 320: Determine a reference gesture frame subsequence, wherein the reference gesture frame subsequence is any one of the gesture frame subsequences, determine the target gesture frame based on the classification information of the reference gesture frame subsequence, and determine the target joint rotation information in a preset joint rotation information database.
[0147] Step 322: Determine the number of supplementary frames based on the number of frames in the reference gesture frame subsequence and the number of frames in the target gesture frame, and generate a joint rotation frame sequence based on the number of supplementary frames and the target joint rotation information.
[0148] Step 324: Generate the target gesture frame sequence based on the target gesture frame and the joint rotation frame sequence.
[0149] One embodiment of this application implements a method for generating gesture frame sequences. It determines the gesture frames to be processed from the video of the action to be processed, reducing the impact of gesture-irrelevant regions on the classification information output and improving the accuracy of the classification information. Then, based on the classification information, at least one initial gesture frame is determined from each of the gesture frames to be processed. Initial labeled gesture frame sequences are constructed by retaining the initial gesture frames with high reliability of the classification results. The initial labeled gesture frame sequences are then adjusted according to the frame information and classification information corresponding to each initial gesture frame, so that the generated reference labeled gesture frame sequence can better reflect the change process of the gesture posture. A target gesture frame sequence is generated based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence. The joint rotation information reflects the gesture action process, improving the smoothness of the gesture frame sequence. The above-described method for generating gesture frame sequences can balance generation efficiency, action accuracy, and smoothness, making it better applicable to different project development and application scenarios.
[0150] Corresponding to the above method embodiments, this application also provides an embodiment of a gesture frame sequence generation apparatus. Figure 4 A schematic diagram of a gesture frame sequence generation device according to an embodiment of this application is shown. Figure 4 As shown, the device includes: The classification module 402 is configured to acquire at least one gesture frame to be processed in the action video to be processed, and determine the classification information corresponding to each gesture frame to be processed. The construction module 404 is configured to determine at least one initial gesture frame from each gesture frame to be processed based on each classification information, and to construct an initial labeled gesture frame sequence based on each initial gesture frame and the classification information corresponding to each initial gesture frame. The adjustment module 406 is configured to adjust the initial labeled gesture frame sequence according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence. The generation module 408 is configured to generate a target gesture frame sequence based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
[0151] Optionally, the classification module 402 is further configured to: Analyze the video of the action to be processed to obtain at least one initial video frame; Identify and extract the gesture region in each initial video frame to obtain at least one gesture frame to be processed.
[0152] Optionally, the classification module 402 is further configured to: Each gesture frame to be processed is input into the gesture classification model to obtain the classification information corresponding to each gesture frame to be processed output by the gesture classification model.
[0153] Optionally, the classification module 402 is further configured to: Obtain an initial video sample set corresponding to at least one preset gesture type, wherein the initial video sample set includes at least one initial video sample corresponding to a video viewpoint; Determine at least one gesture image sample corresponding to each initial video sample, input each gesture image sample into the gesture classification model, and obtain the training classification information output by the gesture classification model; The model loss information is calculated based on the training classification information and the preset gesture type. The model parameters of the gesture classification model are adjusted according to the model loss information, and the gesture classification model is trained until the model training stops.
[0154] Optionally, the classification information includes classification results and classification confidence levels; The building module 404 is further configured as follows: Each gesture frame to be processed with a classification confidence score greater than or equal to the classification confidence threshold is used as the initial gesture frame; Based on each initial gesture frame and the corresponding classification result, construct the first labeled gesture frame sequence; The first labeled gesture frame sequence is updated based on the frame interval between adjacent initial gesture frames and the denoising frame interval threshold in the first labeled gesture frame sequence, and the initial labeled gesture frame sequence is constructed.
[0155] Optionally, the classification information includes classification results and classification confidence, wherein the classification results include initial type and at least one functional type; The adjustment module 406 is further configured to: At least one adjacent initial gesture frame is determined based on the frame information corresponding to each initial gesture frame. Determine the target adjacent initial gesture frame and the target frame interval of the target adjacent initial gesture frame, wherein the target adjacent initial gesture frame is a function type; If the target frame interval is greater than or equal to the frame interval threshold, a reference annotation gesture frame is determined based on the target's adjacent initial gesture frame and the initial gesture frame of the starting type. If the target frame interval is less than the frame interval threshold, a reference labeled gesture frame is determined based on the classification confidence of the target's adjacent initial gesture frames; The reference annotation gesture frame sequence is constructed based on each reference annotation gesture frame.
[0156] Optionally, the adjustment module 406 is further configured to: Insert an initial gesture frame of the starting type into the adjacent initial gesture frames of the target to obtain a reference annotation gesture frame.
[0157] Optionally, the adjustment module 406 is further configured to: The classification confidence scores of each adjacent initial gesture frame are compared to obtain the comparison results; The initial gesture frame to be retained is determined based on the comparison results, and the reference annotation gesture frame is generated.
[0158] Optionally, the generation module 408 is further configured to: Determine a reference gesture frame subsequence, wherein the reference gesture frame subsequence is any one of the gesture frame subsequences; The target gesture frame is determined based on the classification information of the reference gesture frame subsequence, and the target joint rotation information is determined in the preset joint rotation information database. Generate a joint rotation frame sequence based on the reference gesture frame subsequence and the target joint rotation information; The target gesture frame sequence is generated based on the target gesture frame and the joint rotation frame sequence.
[0159] Optionally, the generation module 408 is further configured to: The number of supplementary frames is determined based on the number of frames in the reference gesture frame subsequence and the number of frames in the target gesture frame; A joint rotation frame sequence is generated based on the number of supplementary frames and the target joint rotation information.
[0160] Optionally, the generation module 408 is further configured to: Determine the starting gesture and at least one functional gesture; Obtain motion sampling videos corresponding to each functional gesture, wherein the motion sampling videos include the starting gesture and the target functional gesture, and the target functional gesture is any one of the functional gestures; Each motion sampling video is input into a rotation parameterization model, which outputs preset joint rotation information corresponding to each motion sampling video.
[0161] Optionally, the building module 404 is further configured to: A target gesture frame sequence is generated based on the initial labeled gesture frame sequence and the joint rotation information corresponding to the initial labeled gesture frame sequence.
[0162] One embodiment of this application implements a gesture frame sequence generation device. It determines the gesture frames to be processed from the action video, reducing the impact of gesture-irrelevant regions on the classification information output and improving the accuracy of the classification information. Then, based on the classification information, it determines at least one initial gesture frame from each gesture frame to be processed. It retains the initial gesture frames with high reliability of the classification results to construct an initial labeled gesture frame sequence. The initial labeled gesture frame sequence is then adjusted based on the frame information and classification information corresponding to each initial gesture frame, so that the generated reference labeled gesture frame sequence can better reflect the change process of the gesture posture. A target gesture frame sequence is generated based on each gesture frame sub-sequence and the joint rotation information corresponding to each gesture frame sub-sequence. The joint rotation information reflects the gesture action process, improving the smoothness of the gesture frame sequence. The above-described gesture frame sequence generation device can balance generation efficiency, action accuracy, and smoothness, making it better applicable to different project development and application scenarios.
[0163] The above is an illustrative scheme of a gesture frame sequence generation device according to this embodiment. It should be noted that the technical solution of this gesture frame sequence generation device and the technical solution of the gesture frame sequence generation method described above belong to the same concept. For details not described in detail in the technical solution of the gesture frame sequence generation device, please refer to the description of the technical solution of the gesture frame sequence generation method described above.
[0164] Figure 5 A structural block diagram of a computing device 500 according to an embodiment of this application is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0165] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0166] In one embodiment of this application, the aforementioned components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0167] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0168] The processor 520 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described gesture frame sequence generation method.
[0169] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the gesture frame sequence generation method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the gesture frame sequence generation method described above.
[0170] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the gesture frame sequence generation method described above.
[0171] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the gesture frame sequence generation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the gesture frame sequence generation method described above.
[0172] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the gesture frame sequence generation method described above.
[0173] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the gesture frame sequence generation method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the gesture frame sequence generation method described above.
[0174] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0175] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0176] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0178] The preferred embodiments disclosed above are merely illustrative of this application. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this application. These embodiments are selected and specifically described in this application to better explain the principles and practical applications of this application, thereby enabling those skilled in the art to better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. A method for generating a gesture frame sequence, characterized in that, include: Obtain at least one gesture frame from the video of the action to be processed, and determine the classification information corresponding to each gesture frame. At least one initial gesture frame is determined from each gesture frame to be processed based on the classification information. An initial labeled gesture frame sequence is constructed based on each initial gesture frame and the classification information corresponding to each initial gesture frame. The initial labeled gesture frame sequence is adjusted according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence; The target gesture frame sequence is generated based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
2. The method as described in claim 1, characterized in that, Obtain at least one gesture frame from the video of the action to be processed, including: Analyze the video of the action to be processed to obtain at least one initial video frame; Identify and extract the gesture region in each initial video frame to obtain at least one gesture frame to be processed.
3. The method as described in claim 1, characterized in that, Determine the classification information corresponding to each gesture frame to be processed, including: Each gesture frame to be processed is input into the gesture classification model to obtain the classification information corresponding to each gesture frame to be processed output by the gesture classification model.
4. The method as described in claim 3, characterized in that, The gesture classification model is trained and generated through the following steps: Obtain an initial video sample set corresponding to at least one preset gesture type, wherein the initial video sample set includes at least one initial video sample corresponding to a video viewpoint; Determine at least one gesture image sample corresponding to each initial video sample, input each gesture image sample into the gesture classification model, and obtain the training classification information output by the gesture classification model; The model loss information is calculated based on the training classification information and the preset gesture type. The model parameters of the gesture classification model are adjusted according to the model loss information, and the gesture classification model is trained until the model training stops.
5. The method as described in claim 1, characterized in that, The classification information includes the classification result and the classification confidence level; Based on the classification information, at least one initial gesture frame is determined from each gesture frame to be processed. Based on each initial gesture frame and the classification information corresponding to each initial gesture frame, an initial labeled gesture frame sequence is constructed, including: Each gesture frame to be processed with a classification confidence score greater than or equal to the classification confidence threshold is used as the initial gesture frame; Based on each initial gesture frame and the corresponding classification result, construct the first labeled gesture frame sequence; The first labeled gesture frame sequence is updated based on the frame interval between adjacent initial gesture frames and the denoising frame interval threshold in the first labeled gesture frame sequence, and the initial labeled gesture frame sequence is constructed.
6. The method as described in claim 1, characterized in that, The classification information includes classification results and classification confidence, wherein the classification results include initial type and at least one functional type; The initial labeled gesture frame sequence is adjusted based on the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, including: At least one adjacent initial gesture frame is determined based on the frame information corresponding to each initial gesture frame. Determine the target adjacent initial gesture frame and the target frame interval of the target adjacent initial gesture frame, wherein the target adjacent initial gesture frame is a function type; If the target frame interval is greater than or equal to the frame interval threshold, a reference annotation gesture frame is determined based on the target's adjacent initial gesture frame and the initial gesture frame of the starting type. If the target frame interval is less than the frame interval threshold, a reference labeled gesture frame is determined based on the classification confidence of the target's adjacent initial gesture frames; The reference annotation gesture frame sequence is constructed based on each reference annotation gesture frame.
7. The method as described in claim 6, characterized in that, Determining reference annotation gesture frames based on the target's adjacent initial gesture frames and the initial gesture frame includes: Insert an initial gesture frame of the starting type into the adjacent initial gesture frames of the target to obtain a reference annotation gesture frame.
8. The method as described in claim 6, characterized in that, Determine reference labeled gesture frames based on the classification confidence of the target's adjacent initial gesture frames, including: The classification confidence scores of each adjacent initial gesture frame are compared to obtain the comparison results; The initial gesture frame to be retained is determined based on the comparison results, and the reference annotation gesture frame is generated.
9. The method as described in claim 1, characterized in that, The target gesture frame sequence is generated based on each gesture frame sub-sequence and the joint rotation information corresponding to each gesture frame sub-sequence, including: Determine a reference gesture frame subsequence, wherein the reference gesture frame subsequence is any one of the gesture frame subsequences; The target gesture frame is determined based on the classification information of the reference gesture frame subsequence, and the target joint rotation information is determined in the preset joint rotation information database. Generate a joint rotation frame sequence based on the reference gesture frame subsequence and the target joint rotation information; The target gesture frame sequence is generated based on the target gesture frame and the joint rotation frame sequence.
10. The method as described in claim 9, characterized in that, Generate a joint rotation frame sequence based on the reference gesture frame subsequence and the target joint rotation information, including: The number of supplementary frames is determined based on the number of frames in the reference gesture frame subsequence and the number of frames in the target gesture frame; A joint rotation frame sequence is generated based on the number of supplementary frames and the target joint rotation information.
11. The method as described in claim 9, characterized in that, The preset joint rotation information database is constructed through the following steps: Determine the starting gesture and at least one functional gesture; Obtain motion sampling videos corresponding to each functional gesture, wherein the motion sampling videos include the starting gesture and the target functional gesture, and the target functional gesture is any one of the functional gestures; Each motion sampling video is input into a rotation parameterization model, which outputs preset joint rotation information corresponding to each motion sampling video.
12. The method as described in claim 1, characterized in that, After constructing the initial labeled gesture frame sequence, the method further includes: A target gesture frame sequence is generated based on the initial labeled gesture frame sequence and the joint rotation information corresponding to the initial labeled gesture frame sequence.
13. A gesture frame sequence generation device, characterized in that, include: The classification module is configured to acquire at least one gesture frame to be processed in the action video to be processed, and determine the classification information corresponding to each gesture frame to be processed. The construction module is configured to determine at least one initial gesture frame from each gesture frame to be processed based on each classification information, and to construct an initial labeled gesture frame sequence based on each initial gesture frame and the classification information corresponding to each initial gesture frame. The adjustment module is configured to adjust the initial labeled gesture frame sequence according to the frame information and classification information corresponding to each initial gesture frame to obtain a reference labeled gesture frame sequence, wherein the reference labeled gesture frame sequence includes at least one gesture frame subsequence; The generation module is configured to generate a target gesture frame sequence based on each gesture frame subsequence and the joint rotation information corresponding to each gesture frame subsequence.
14. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 12.