Portrait posture control method, device, computer equipment and storage medium

By decoupling expression and posture control, using posture and expression feature parameters to generate portrait posture control, the problem of poor generation effect in the existing technology is solved, and high-precision portrait posture control is realized, which is suitable for high-demand fields such as film and television production.

CN120103877BActive Publication Date: 2025-08-12GUANGZHOU QUWAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510579223.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-12
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing portrait attitude control methods have poor generation results in high-quality scenarios, which are difficult to meet the high-standard needs of portrait attitude control for film and television video characters, and cannot be applied to application scenarios with real-time and high generation quality requirements, and have low accuracy.

Method used

By obtaining the original portrait and reference video clips, the pre-trained portrait posture control model extracts the posture characteristics and expression features, calculates the posture control parameters and expression control parameters respectively, realizes the decoupling control of expression and attitude, uses the global sample video clip to calculate the expression control parameters, and local reference video clips to calculate the attitude control parameters.

Benefits of technology

It improves the accuracy of portrait posture control, can generate high-fidelity and natural portrait action videos, adapt to diverse application needs, and meet technical support in high-demand fields such as film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103877B_ABST
    Figure CN120103877B_ABST
Patent Text Reader

Abstract

This application relates to a portrait pose control method, apparatus, computer device, and storage medium. The method includes: obtaining an original first portrait whose portrait pose is to be controlled, and a reference video clip for performing portrait pose control on the original first portrait; inputting the original first portrait and the reference video clip into a pre-trained portrait pose control model, obtaining pose features of each second portrait in the reference video clip using the portrait pose control model, and obtaining pose control parameters based on the pose features; obtaining expression features of each sample portrait in a set of sample video clips, and obtaining expression control parameters based on the expression features; the sample video clip set being a set of sample video clips used to train the portrait pose control model; and controlling the original first portrait using the expression control parameters and the pose control parameters to generate a target first portrait after portrait pose control. This method can effectively improve the accuracy of portrait pose control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a portrait posture control method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0002] With the development of artificial intelligence technology, a technology that uses artificial intelligence to achieve portrait posture control has emerged. This technology allows users to easily adjust and drive the posture movements of the portraits of people in the video. With the help of this technology, there is no need for repeated shooting or re-production, and video content with different posture movements can be quickly generated in batches, thereby significantly improving creative efficiency and optimizing the production process. Therefore, its application in the video field is becoming increasingly widespread.

[0003] However, the currently provided portrait pose control methods have poor generation effects in high-quality scenes, which makes it difficult to meet the high-standard requirements of film-level video portrait pose control. They cannot also be applied to application scenarios that require real-time performance and high generation quality. Therefore, the existing portrait pose control has low accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a portrait posture control method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the portrait posture control accuracy in order to address the above technical problems.

[0005] In a first aspect, the present application provides a portrait gesture control method, comprising:

[0006] Acquire an original first portrait of a to-be-controlled portrait posture, and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait;

[0007] Inputting the original first portrait and the reference video clip into a pre-trained portrait pose control model, obtaining pose features of each second portrait in the reference video clip through the portrait pose control model, and obtaining pose control parameters based on the pose features of each second portrait;

[0008] Obtaining expression features of each sample portrait in a sample video clip set, and obtaining expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used for training the portrait posture control model, and each sample video clip carries multiple frames of sample portraits;

[0009] The original first portrait is controlled by using the expression control parameters and the posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait.

[0010] In one embodiment, the use of the expression control parameters and the posture control parameters to control the original first portrait to generate a target first portrait after portrait posture control includes: obtaining the expression characteristics of the original first portrait and the posture characteristics of the original first portrait; standardizing the expression characteristics of the original first portrait and the posture characteristics of the original first portrait to obtain standardized expression characteristics and standardized posture characteristics; adjusting the standardized expression characteristics using the expression control parameters to obtain the expression characteristics of the target first portrait, and adjusting the standardized posture characteristics using the posture control parameters to obtain the posture characteristics of the target first portrait; generating the target first portrait based on the expression characteristics of the target first portrait and the posture characteristics of the target first portrait.

[0011] In one embodiment, the expression control parameters include: an expression feature mean and an expression feature standard deviation obtained based on the expression features of each of the sample portraits; the posture control parameters include: a posture feature mean and a posture feature standard deviation obtained based on the posture features of each of the second portraits; using the expression control parameters to adjust the standardized expression features to obtain the expression features of the target first portrait, and using the posture control parameters to adjust the standardized posture features to obtain the posture features of the target first portrait, include: mapping the expression feature mean and the expression feature standard deviation to the standardized expression features to obtain the expression features of the target first portrait; mapping the posture feature mean and the posture feature standard deviation to the standardized posture features to obtain the posture features of the target first portrait.

[0012] In one embodiment, the portrait posture control model includes an encoder and a decoder; obtaining the expression features of the original first portrait and the posture features of the original first portrait includes: mapping the original first portrait to a three-dimensional latent space through the encoder, generating three-dimensional key points in the three-dimensional latent space, and encoding and generating the expression features of the original first portrait and the posture features of the original first portrait according to the position coordinates of the three-dimensional key points; generating the target first portrait according to the expression features of the target first portrait and the posture features of the target first portrait includes: inputting the expression features of the target first portrait and the posture features of the target first portrait into the decoder, upsampling and decoding the expression features of the target first portrait and the posture features of the target first portrait through the decoder, and outputting the target first portrait.

[0013] In one embodiment, the portrait posture control model includes an encoder and a decoder; the portrait posture control model is trained by the following steps: each of the sample video clips in the sample video clip set is input into the portrait posture control model to be trained, and the expression features and posture features of the multiple frames of sample portraits contained in each of the sample video clips are obtained through the encoder of the portrait posture control model; the expression features and posture features of the multiple frames of sample portraits are input into the decoder of the portrait posture control model, and the reconstructed video clips corresponding to each of the sample video clips are generated through the decoder; based on the difference between the reconstructed video clips and the sample video clips, the portrait posture control model is trained to obtain the pre-trained portrait posture control model.

[0014] In one embodiment, the step of inputting the expression features and posture features of the multiple frames of sample portraits into the decoder of the portrait posture control model, and generating reconstructed video segments corresponding to each of the sample video segments through the decoder, includes: inputting the expression features and posture features of each frame of sample portraits contained in the current sample video segment into the decoder, and obtaining reconstructed portraits of each frame corresponding to the current sample video segment through the decoder; the current sample video segment is any one of the sample video segments in the sample video segment set; and combining the reconstructed portraits of each frame corresponding to the current sample video segment to obtain a reconstructed video segment corresponding to the current sample video segment.

[0015] In a second aspect, the present application further provides a portrait gesture control device, comprising:

[0016] an original portrait acquisition module, configured to acquire an original first portrait of a portrait posture to be controlled, and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait;

[0017] a posture parameter acquisition module, configured to input the original first portrait and the reference video clip into a pre-trained portrait posture control model, obtain posture features of each second portrait in the reference video clip through the portrait posture control model, and obtain posture control parameters based on the posture features of each second portrait;

[0018] an expression parameter acquisition module, configured to acquire expression features of each sample portrait in a sample video clip set, and acquire expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used to train the portrait gesture control model, each sample video clip carrying multiple frames of sample portraits;

[0019] The target portrait generation module is used to control the original first portrait using the expression control parameters and the posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait.

[0020] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in any one of the embodiments of the first aspect when executing the computer program.

[0021] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.

[0022] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.

[0023] The above-mentioned portrait posture control method, device, computer equipment, storage medium and computer program product obtain an original first portrait of a portrait posture to be controlled, and a reference video clip for portrait posture control of the original first portrait; the reference video clip carries multiple frames of second portraits; the original first portrait and the reference video clip are input into a pre-trained portrait posture control model, the posture features of each second portrait in the reference video clip are obtained through the portrait posture control model, and posture control parameters are obtained according to the posture features of each second portrait; the expression features of each sample portrait in a sample video clip set are obtained, and expression control parameters are obtained according to the expression features of each sample portrait; the sample video clip set is a set of sample video clips used for training the portrait posture control model, and each sample video clip carries multiple frames of sample portraits; the original first portrait is controlled using the expression control parameters and the posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait. The present application obtains an original first portrait whose portrait posture needs to be controlled, and a reference video clip carrying multiple frames of a second portrait for portrait posture control, thereby inputting the original first portrait and the reference video clip into a pre-trained portrait posture control model. The model can obtain posture control parameters based on the reference video clip, and then obtain expression control parameters based on the sample video clip used to train the portrait posture control model, thereby generating a target first portrait using the expression control parameters and the posture control parameters. In this way, the decoupling of expression control and posture control is achieved. At the same time, expression control uses global sample video clips to calculate control parameters, while posture control only uses reference video clips to calculate control parameters. In this way, posture control can take into account the differences between different identities, and expression control can be better generalized to unseen expression types. Therefore, the portrait control achieved in this way can effectively improve the accuracy of portrait posture control. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 is a flow chart of a portrait gesture control method according to one embodiment;

[0026] Figure 2 A schematic diagram of a process for generating a first portrait of a target in one embodiment;

[0027] Figure 3 1 is a schematic diagram of a process for training a portrait posture control model in one embodiment;

[0028] Figure 4 is a structural block diagram of a portrait gesture control device in one embodiment;

[0029] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0031] In one embodiment, Figure 1 As shown, a portrait posture control method is provided. This embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0032] Step S101 : obtaining an original first portrait of a to-be-controlled portrait posture and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait.

[0033] The original first portrait refers to the original portrait whose portrait posture needs to be controlled, that is, the original portrait whose portrait posture needs to be changed. The number of the original first portraits can be single or multiple, and the reference video clip is a video clip used to implement portrait posture control of the original first portrait. The video clip can be composed of multiple video frames, and the video frame can carry the second portrait, so the reference video clip carries multiple frames of the second portrait.

[0034] Specifically, when controlling the portrait posture of the original first portrait, the original first portrait whose portrait posture needs to be controlled and a video clip for controlling the portrait posture of the original first portrait, i.e., a reference video clip, can be obtained first. The reference video clip can carry multiple frames of the second portrait.

[0035] Step S102 : Input the original first portrait and the reference video clip into a pre-trained portrait posture control model, obtain posture features of each second portrait in the reference video clip through the portrait posture control model, and obtain posture control parameters according to the posture features of each second portrait.

[0036] The portrait posture control model is a pre-trained neural network model for posture control of the portrait, and the posture control parameter refers to the control parameter used to control the posture, which can be calculated based on the posture features of each second portrait in the reference video clip.

[0037] Specifically, after the server obtains the original first portrait and the reference video clip, it can input the original first portrait and the reference video clip into a pre-trained portrait posture control model for portrait posture control of the original first portrait. Through this portrait posture control model, the posture features of each second portrait can be first extracted from the reference video clip. The posture features of each second portrait are further used to generate posture control parameters for posture control. Because the posture control parameters are calculated based on the posture features of the input reference video clip, they can fully account for the differences between different video clips. This method can better capture the changing patterns of individual head posture without being affected by other video clips.

[0038] Step S103, obtaining the expression features of each sample portrait in the sample video clip set, and obtaining expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used for training the portrait posture control model, and each sample video clip carries multiple frames of sample portraits.

[0039] A sample video clip set refers to a set of sample video clips. Sample video clips refer to video clips used to train the portrait pose control model. Sample portraits are portraits contained in the sample video clips. Since each sample video clip is composed of multiple video image frames, each sample video clip can contain multiple frames of sample portraits. Expression control parameters are control parameters used to control expressions. Unlike pose control parameters, in this embodiment, expression control parameters are not derived based on the expression features of the portraits in the reference video clips, but rather are derived from the expression features of the portraits contained in each sample video clip in the sample video clip set. That is, they are derived through global expression feature extraction.

[0040] Specifically, the server can also use the facial features of each sample portrait included in the sample video clip set used to train the portrait pose control model to generate expression control parameters for controlling facial expressions. Because these expression control parameters are calculated based on the facial features of each sample portrait in the global sample video clip set, they enable the model to better generalize to unseen expression types.

[0041] Step S104 , controlling the original first portrait using expression control parameters and posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait.

[0042] The target first portrait refers to a portrait generated by controlling the original first portrait using expression control parameters and posture control parameters. The portrait can be generated by deforming the original first portrait using expression control parameters and posture control parameters. The expression control parameters are mainly used to control the expression of the generated target first portrait, while the posture control parameters are mainly used to control the posture of the generated target first portrait.

[0043] Specifically, after obtaining the expression control parameters and the posture control parameters, the server can use the expression control parameters to adjust the expression of the original first portrait, and use the posture control parameters to adjust the posture of the original first portrait, thereby generating an adjusted first portrait, that is, generating a target first portrait after portrait posture control.

[0044] In the above-mentioned portrait posture control method, an original first portrait of a portrait posture to be controlled and a reference video clip for performing portrait posture control on the original first portrait are obtained through a server; the reference video clip carries multiple frames of second portraits; the original first portrait and the reference video clip are input into a pre-trained portrait posture control model, and the posture features of each second portrait in the reference video clip are obtained through the portrait posture control model, and posture control parameters are obtained based on the posture features of each second portrait; the expression features of each sample portrait in a sample video clip set are obtained, and expression control parameters are obtained based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used for training the portrait posture control model, and each sample video clip carries multiple frames of sample portraits; the expression control parameters and posture control parameters are used to control the original first portrait to generate a target first portrait after portrait posture control; wherein, the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait. The present application obtains an original first portrait whose portrait posture needs to be controlled, and a reference video clip carrying multiple frames of a second portrait for portrait posture control, thereby inputting the original first portrait and the reference video clip into a pre-trained portrait posture control model. The model can obtain posture control parameters based on the reference video clip, and then obtain expression control parameters based on the sample video clip used to train the portrait posture control model, thereby generating a target first portrait using the expression control parameters and the posture control parameters. In this way, the decoupling of expression control and posture control is achieved. At the same time, expression control uses global sample video clips to calculate control parameters, while posture control only uses reference video clips to calculate control parameters. In this way, posture control can take into account the differences between different identities, and expression control can be better generalized to unseen expression types. Therefore, the portrait control achieved in this way can effectively improve the accuracy of portrait posture control.

[0045] In one embodiment, Figure 2 As shown, step S104 may further include:

[0046] Step S201, acquiring the expression features of the original first portrait and the posture features of the original first portrait;

[0047] Step S202 : performing standardization processing on the expression features and the posture features of the original first portrait to obtain standardized expression features and standardized posture features.

[0048] The standardized expression feature refers to a feature generated by standardizing the expression feature of the original first portrait, and the standardized posture feature refers to a feature generated by standardizing the posture feature of the original first portrait.

[0049] Specifically, after the server inputs the original first portrait into the pre-trained portrait posture control model, it can first extract the expression features and posture features of the original first portrait, and then perform standardization based on the extracted expression features and posture features of the original first portrait. For example, the mean and standard deviation of the expression features of the original first portrait, as well as the mean and standard deviation of the posture features of the original first portrait, can be calculated respectively, and then the expression features of the original first portrait can be standardized using the mean and standard deviation of the expression features, and the posture features of the original first portrait can be standardized using the mean and standard deviation of the posture features, thereby obtaining standardized expression features and standardized posture features.

[0050] Step S203, adjusting the standardized expression features using the expression control parameters to obtain expression features of the target first portrait, and adjusting the standardized posture features using the posture control parameters to obtain posture features of the target first portrait;

[0051] Step S204 : generating a first target portrait according to the expression features of the first target portrait and the posture features of the first target portrait.

[0052] Afterwards, the server can use the expression control parameters to adjust the standardized expression features to obtain the expression features of the target first portrait, and use the posture control parameters to adjust the standardized posture features to obtain the posture features of the target first portrait. Finally, the expression features and posture features of the target first portrait can be used to generate the target first portrait after the original first portrait is deformed.

[0053] In this embodiment, the expression features and posture features of the original first portrait can also be extracted through the portrait posture control model, and the above features are standardized. After obtaining the standardized features, the above standardized features are adjusted using corresponding control parameters to obtain the expression features and posture features of the target first portrait, thereby generating the target first portrait. In this way, the accuracy of generating the target first portrait can be improved.

[0054] Furthermore, the expression control parameters include: the expression feature mean and the expression feature standard deviation obtained based on the expression features of each sample portrait, and the posture control parameters include: the posture feature mean and the posture feature standard deviation obtained based on the posture features of each second portrait; step S203 may further include: mapping the expression feature mean and the expression feature standard deviation to standardized expression features to obtain the expression features of the target first portrait; mapping the posture feature mean and the posture feature standard deviation to standardized posture features to obtain the posture features of the target first portrait.

[0055] In this embodiment, the expression control parameter for controlling facial expressions may be composed of two parts: the mean and standard deviation calculated based on the facial expression features of each sample portrait, i.e., the expression feature mean and the expression feature standard deviation. Similarly, the posture control parameter for controlling posture may also be composed of two parts: the mean and standard deviation calculated based on the posture features of each second portrait, i.e., the posture feature mean and the posture feature standard deviation.

[0056] Specifically, after obtaining the posture features of each second portrait, the server can also calculate the mean and standard deviation based on the above posture features as posture control parameters, and after obtaining the expression features of each sample portrait, it can also calculate the mean and standard deviation based on the above expression features as expression control parameters.

[0057] Afterwards, the above-mentioned expression control parameters can be used to adjust the standardized expression features of the original first portrait, and the above-mentioned posture control parameters can be used to adjust the standardized posture features of the original first portrait, that is, the expression feature mean and the expression feature standard deviation can be mapped to the standardized expression feature to obtain the expression features of the target first portrait, that is, the expression feature mean and the standard deviation of the global sample portrait are applied to the standardized expression features of the original first portrait to obtain the expression features of the target first portrait. Similarly, for the posture features, the posture feature mean and the posture feature standard deviation can be mapped to the standardized posture features to obtain the posture features of the target first portrait, that is, the posture feature mean and the standard deviation of the local second portrait in the reference video clip are applied to the standardized posture features of the original first portrait to obtain the posture features of the target first portrait.

[0058] In this embodiment, the mean and standard deviation can be calculated based on the expression features of each sample portrait as expression control parameters, and the above mean and standard deviation can be mapped to the standardized expression features to achieve the generation of expression features of the target first portrait, and the mean and standard deviation can be calculated based on the posture features of each second portrait as posture control parameters, and the above mean and standard deviation can be mapped to the standardized posture features to achieve the generation of posture features of the target first portrait. In this way, the expression features can be standardized using global statistics, and the posture features can be standardized by calculating the mean and standard deviation using independent video clips. In this way, the expression can be operated under a global distribution, and the posture can be migrated according to different character portrait style habits, so the accuracy of the generated posture can be ensured.

[0059] In addition, the portrait posture control model may include an encoder and a decoder; step S201 may further include: mapping the original first portrait to a three-dimensional latent space through the encoder, and generating three-dimensional key points in the three-dimensional latent space, and encoding and generating the expression features and posture features of the original first portrait according to the position coordinates of the three-dimensional key points; step S204 may further include: inputting the expression features of the target first portrait and the posture features of the target first portrait into the decoder, upsampling and decoding the expression features and posture features of the target first portrait through the decoder, and outputting the target first portrait.

[0060] In this embodiment, the portrait pose control model may include an encoder and a decoder, wherein the encoder is primarily used to extract portrait image features, while the decoder is used to reconstruct the features into a portrait image. Furthermore, to enable real-time, accurate, and flexible control of the portrait pose of a video or image, an editable implicit spatial operation on the image may be configured to extract portrait features and output the portrait image.

[0061] Specifically, after the original first portrait is input into a pre-trained portrait pose control model, the original first portrait can be mapped to a three-dimensional latent space through the encoder of the portrait pose control model. Then, a set of three-dimensional key points can be generated in the latent space, so as to obtain the expression features and posture features of the original first portrait based on the coordinates of the three-dimensional key points, realizing the encoding and decoupling of features to control the position deformation of features.

[0062] Afterwards, the expression features and posture features of the original first portrait can be deformed in the three-dimensional latent space. For example, the expression features and posture features of the original first portrait can be standardized in the three-dimensional latent space first, and then the posture control parameters can be used to adjust the standardized expression features and posture features in the three-dimensional latent space to achieve feature deformation, thereby obtaining the expression features and posture features of the target first portrait.

[0063] Finally, the server may also input the expression features and posture features of the target first portrait into the decoder in the portrait posture control model, and upsample and decode the expression features and posture features of the target first portrait through the decoder to finally output the target first portrait.

[0064] In this embodiment, a three-dimensional latent space can also be constructed through the encoder in the portrait posture control model, so as to map the original first portrait to the three-dimensional latent space, and extract the expression features and posture features of the first portrait from the three-dimensional latent space. Then, the decoder in the portrait posture control model can be used to generate the target first portrait output according to the adjusted expression features and posture features in the three-dimensional latent space. In this way, the portrait posture can be controlled in real time, accurately and flexibly.

[0065] In one embodiment, the portrait gesture control model includes an encoder and a decoder; Figure 3 As shown in Figure 2, the portrait pose control model is trained through the following steps:

[0066] Step S301: input each sample video clip in the sample video clip set into the portrait posture control model to be trained, and obtain the expression features and posture features of multiple frames of sample portraits contained in each sample video clip through the encoder of the portrait posture control model.

[0067] The portrait pose control model to be trained refers to an untrained portrait pose control model. In this embodiment, the portrait pose control model may include an encoder and a decoder, wherein the encoder is primarily used to extract portrait features, and the decoder is primarily used to output a corresponding portrait based on the portrait features. Specifically, when training the portrait pose control model, sample video clips included in a pre-collected set of sample video clips may be input into the decoder of the portrait pose control model. The encoder then extracts the expression features and pose features of multiple frames of sample portraits contained in each sample video clip.

[0068] Step S302: inputting the expression features and posture features of the multiple frames of sample portraits into a decoder of a portrait posture control model, and generating a reconstructed video segment corresponding to each sample video segment through the decoder;

[0069] Step S303 : training a portrait pose control model based on the difference between the reconstructed video segment and the sample video segment to obtain a pre-trained portrait pose control model.

[0070] Since the portrait posture control model provided in this embodiment does not require additional learning for processing expression features and posture features, the model training process is mainly used to learn the process of extracting portrait features and learning the process of reconstructing portraits using portrait features, and the reconstructed video clips are video clips generated after the portraits are reconstructed using portrait features. Therefore, in this embodiment, after obtaining the expression features and posture features of multiple frames of sample portraits, the expression features and posture features of the above multiple frames of sample portraits can be input into the decoder of the portrait posture control model, thereby realizing portrait reconstruction through the decoder to generate reconstructed video clips corresponding to each sample video clip. Based on the differences between the reconstructed video clips and the sample video clips, the portrait posture control model is trained to obtain a pre-trained portrait posture control model.

[0071] In this embodiment, the sample video clips in the sample video clip set can also be input into the portrait posture control model, and the encoder of the portrait posture control model extracts the expression features and posture features of the multiple frames of sample portraits contained in the sample video clips, and uses the decoder to reconstruct the expression features and posture features. After obtaining the reconstructed video clip, the portrait posture control model is completed by using the difference between the reconstructed video clip and the sample video clip. In this way, the model only needs to learn the process of portrait feature extraction and portrait reconstruction, without having to learn the process of using control parameters to control the expression features and posture features, thereby improving the efficiency of portrait posture control model training.

[0072] Furthermore, step S302 may further include: inputting the expression features and posture features of each frame of sample portrait contained in the current sample video clip into the decoder, and obtaining the reconstructed portraits of each frame corresponding to the current sample video clip through the decoder; the current sample video clip is any one of the sample video clips in the sample video clip set; and combining the reconstructed portraits of each frame corresponding to the current sample video clip to obtain a reconstructed video clip corresponding to the current sample video clip.

[0073] The current sample video clip refers to any one of the multiple sample video clips in the sample video clip set, and the current sample video clip may also be composed of multiple frames of sample portraits. For example, the sample video clip set may include video clip A, video clip B, and video clip C. Taking video clip A as the current sample video clip as an example, the video clip may include multiple frames of sample portraits, namely, sample portrait A1, sample portrait A2, and sample portrait A3. When the expression features and posture features of the sample portraits of each frame in the current sample video clip are input into the decoder, that is, the expression features and posture features of the sample portrait A1, the sample portrait A2, and the sample portrait A3 are input into the decoder respectively, so that the decoder uses the expression features and posture features of the sample portraits of each frame to generate a reconstructed portrait, that is, the expression features and posture features of the sample portrait A1 are used to generate the reconstructed portrait A1, the expression features and posture features of the sample portrait A2 are used to generate the reconstructed portrait A2, and the expression features and posture features of the sample portrait A3 are used to generate the reconstructed portrait A3.

[0074] After obtaining the reconstructed image for each frame corresponding to the current sample video segment, the reconstructed portraits of each frame can be combined. That is, reconstructed portraits A1, A2, and A3 can be mechanically combined to generate reconstructed video segment A, which serves as the reconstructed video segment corresponding to the current sample video segment. In this way, a reconstructed video segment corresponding to each sample video segment can be obtained.

[0075] In this embodiment, the expression features and posture features of each frame of the sample portrait contained in the current sample video clip can also be input into the decoder to obtain the reconstructed portrait of each frame, and then the reconstructed portraits of each frame are combined to obtain the reconstructed video clip. In this way, the efficiency of video clip reconstruction can be improved.

[0076] In one embodiment, a portrait pose control method based on adaptive standardization is also provided, which can achieve precise control of the portrait pose of a person in different scenes. This method can not only generate high-fidelity and natural portrait motion videos, but also effectively adapt to diverse application requirements, thereby providing higher-quality technical support for film and television production and other demanding fields. The method can be implemented through the following steps:

[0077] 1. Image deformation network and generator:

[0078] In order to be able to control the portrait pose of a video or image in real time, accurately, and flexibly, this embodiment operates on an editable implicit image space. Therefore, the following image deformation network design is mainly carried out:

[0079] (1) Image 3D latent space mapping: Since RGB images have strong coupling and low-order high-frequency features, a model based on a 3D convolutional neural network is designed to map 2D images to a 3D latent space, achieve feature encoding and decoupling, and provide a basis for subsequent spatial deformation.

[0080] (2) 3D latent space keypoints generation: A set of approximately 20 3D keypoints (with coordinates of x, y, and z) is generated in the latent space to control the positional deformation of the feature. Each keypoint defines its influence range and degree through Gaussian distribution, and the surrounding features are weighted and summed to finally generate the deformation feature.

[0081] (3) Spatial Local Deformation: To support independent editing of global and local portrait poses, the system ensures that different keypoints affect only specific regions (e.g., mouth opening and closing does not affect eyes). Independent deformation between regions is achieved by calculating reconstruction losses to constrain the control range of keypoints.

[0082] (4) 3D latent space feature decoding: The deformed 3D latent space features are upsampled and decoded through the generator network (based on the SPADE structure), and finally an RGB image with a user-defined posture is output.

[0083] 2. Adaptive normalized attitude control:

[0084] The above model constructs a latent space representing the image portrait. Therefore, pose vectors in the latent space can be directly transferred, linearly combined, and reorganized. This is then passed through a decoding network to generate a portrait with a modified pose. However, this method represents facial attributes (such as expression and pose) as a coupled latent space, making it difficult to independently control these attributes. For example, the generated result may exhibit pose jitter, lip sync, or strange expressions. The generated video may also contain temporal discontinuities, such as jumps in pose or sudden changes in expression. To address these issues, this embodiment introduces an adaptive pose normalization method to effectively address this coupled transfer problem.

[0085] (1) Adaptive normalization ensures that the expression features and posture features remain independent during the generation process by normalizing them separately. This decoupling mechanism can significantly reduce problems such as posture jitter, lip sync, and strange expressions.

[0086] (2) For facial expression features, calculate the global mean and standard deviation:

[0087]

[0088]

[0089] in, represents the number of frames of the i-th video clip, M represents the total number of training samples, represents the facial expression feature of the jth frame in the i-th video clip, represents the mean value of the facial expression feature, Indicates the standard deviation of the expression feature.

[0090] For pose features, the local mean and standard deviation are calculated for each video segment separately:

[0091]

[0092]

[0093] in, represents the number of frames of the i-th video clip, represents the posture feature of the jth frame in the i-th video clip, represents the mean of the posture features, Indicates the standard deviation of the posture feature.

[0094] By using the above formula to perform corresponding adaptive normalization operations on the expression vector and posture vector in the latent space, it can be ensured that the operation is performed under the global distribution of expression. However, posture is different. If it is also migrated and driven by the global distribution, problems such as portrait jitter and deformation are prone to occur. Adaptive normalization migrates according to different character portrait style habits (the mean and variance are a good representation of each person's action posture), thereby ensuring that the generated posture is continuous and normal.

[0095] Moreover, this adaptive normalization method does not require a training process like mainstream methods such as adaptNorm. With an inference speed of > 25fps, it can flexibly and quickly achieve the migration and control of different character posture habits.

[0096] (3) Effect: The adaptive normalization module significantly improves the temporal continuity and visual quality of the generated video through precise feature decoupling. The module also performs well in reducing posture jitter.

[0097] 3. Model training:

[0098] Since the adaptive normalization operation does not require external model training, the training of the model mainly occurs on the construction of the 2D portrait latent space.

[0099] (1) Model architecture:

[0100] Image 3D latent space mapping module, 3D latent space keypoints generation module, spatial local deformation module, and 3D latent space feature decoding module.

[0101] (2) Model training:

[0102] The model was trained for 100 epochs using portrait video data, and the final model converged well and performed well in quantitative indicators.

[0103] (3) Model conversion:

[0104] The model is converted and packaged using an open source quantization framework to meet the video editing needs of ordinary users.

[0105] In this embodiment, the ability to decouple facial expressions from head posture can be enhanced through a targeted normalization strategy, combined with global and local statistics and transfer loss. This technology improves the accuracy and flexibility of model generation. Specifically, adaptive normalization achieves a finer decoupling by distinguishing and processing facial expressions and portrait posture features. This design solves the problem of easy coupling between facial expressions and head posture in traditional methods and improves the flexibility of model generation. The head posture features are calculated independently by identity using the mean and standard deviation, which fully considers the differences between different video clips. This method can better capture the changing patterns of individual head postures without being affected by other identities. In addition, global statistics can be used to normalize facial expression features, so that the model can better generalize to unseen expression types. At the same time, the application of local statistics also improves the robustness of the model in terms of head posture.

[0106] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0107] Based on the same inventive concept, embodiments of the present application further provide a portrait posture control device for implementing the aforementioned portrait posture control method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more embodiments of the portrait posture control device provided below can be found in the aforementioned limitations of the portrait posture control method and will not be further elaborated here.

[0108] In one embodiment, Figure 4 As shown, a portrait posture control device is provided, comprising: an original portrait acquisition module 401, a posture parameter acquisition module 402, an expression parameter acquisition module 403 and a target portrait generation module 404, wherein:

[0109] The original portrait acquisition module 401 is used to acquire an original first portrait of a portrait posture to be controlled, and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait;

[0110] a pose parameter acquisition module 402 for inputting the original first portrait and the reference video clip into a pre-trained portrait pose control model, acquiring pose features of each second portrait in the reference video clip using the portrait pose control model, and acquiring pose control parameters based on the pose features of each second portrait;

[0111] Expression parameter acquisition module 403 is used to obtain expression features of each sample portrait in the sample video clip set, and obtain expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used for training the portrait posture control model, and each sample video clip carries multiple frames of sample portraits;

[0112] The target portrait generation module 404 is used to control the original first portrait using expression control parameters and posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait.

[0113] In one embodiment, the target portrait generation module 404 is further used to obtain the expression features of the original first portrait and the posture features of the original first portrait; standardize the expression features of the original first portrait and the posture features of the original first portrait to obtain standardized expression features and standardized posture features; adjust the standardized expression features using the expression control parameters to obtain the expression features of the target first portrait, and adjust the standardized posture features using the posture control parameters to obtain the posture features of the target first portrait; and generate the target first portrait based on the expression features of the target first portrait and the posture features of the target first portrait.

[0114] In one embodiment, the expression control parameters include: the expression feature mean and the expression feature standard deviation obtained based on the expression features of each of the sample portraits; the posture control parameters include: the posture feature mean and the posture feature standard deviation obtained based on the posture features of each of the second portraits; the target portrait generation module 404 is further used to map the expression feature mean and the expression feature standard deviation to the standardized expression features to obtain the expression features of the target first portrait; and map the posture feature mean and the posture feature standard deviation to the standardized posture features to obtain the posture features of the target first portrait.

[0115] In one embodiment, the portrait posture control model includes an encoder and a decoder; the target portrait generation module 404 is further used to map the original first portrait to a three-dimensional latent space through the encoder, and generate three-dimensional key points in the three-dimensional latent space, and encode and generate the expression features and posture features of the original first portrait according to the position coordinates of the three-dimensional key points; and is used to input the expression features and posture features of the target first portrait into the decoder, and upsample and decode the expression features and posture features of the target first portrait through the decoder to output the target first portrait.

[0116] In one embodiment, the portrait posture control model includes an encoder and a decoder; the portrait posture control device includes: a control model training module, used to input each of the sample video clips in the sample video clip set into the portrait posture control model to be trained, and obtain the expression features and posture features of multiple frames of sample portraits contained in each of the sample video clips through the encoder of the portrait posture control model; input the expression features and posture features of the multiple frames of sample portraits into the decoder of the portrait posture control model, and generate reconstructed video clips corresponding to each of the sample video clips through the decoder; based on the difference between the reconstructed video clips and the sample video clips, train the portrait posture control model to obtain the pre-trained portrait posture control model.

[0117] In one embodiment, the control model training module is further used to input the expression features and posture features of each frame of sample portrait contained in the current sample video clip into the decoder, and obtain the reconstructed portraits of each frame corresponding to the current sample video clip through the decoder; the current sample video clip is any one of the sample video clips in the sample video clip set; the reconstructed portraits of each frame corresponding to the current sample video clip are combined to obtain the reconstructed video clip corresponding to the current sample video clip.

[0118] Each module in the portrait gesture control device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0119] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store portrait data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a portrait posture control method is implemented.

[0120] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0121] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0123] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0125] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0126] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0127] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A portrait posture control method, characterized in that: The method comprises: Acquire an original first portrait of a to-be-controlled portrait posture, and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait; Inputting the original first portrait and the reference video clip into a pre-trained portrait pose control model, obtaining pose features of each second portrait in the reference video clip using the portrait pose control model, and obtaining pose control parameters based on the pose features of each second portrait; the pose control parameters comprising: a pose feature mean and a pose feature standard deviation obtained based on the pose features of each second portrait; Obtaining expression features of each sample portrait in a sample video clip set, and obtaining expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used to train the portrait posture control model, each sample video clip carrying multiple frames of sample portraits; the expression control parameters include: an expression feature mean and an expression feature standard deviation obtained based on the expression features of each sample portrait; Controlling the original first portrait using the expression control parameters and the posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait; The step of controlling the original first portrait by using the expression control parameter and the posture control parameter to generate a target first portrait after the portrait posture is controlled includes: Acquiring expression features of the original first portrait and posture features of the original first portrait; performing standardization processing on the expression features of the original first portrait and the posture features of the original first portrait to obtain standardized expression features and standardized posture features; The standardized expression feature is adjusted using the expression control parameter to obtain the expression feature of the first target portrait, and the standardized posture feature is adjusted using the posture control parameter to obtain the posture feature of the first target portrait; including: mapping the expression feature mean and the expression feature standard deviation to the standardized expression feature to obtain the expression feature of the first target portrait; mapping the posture feature mean and the posture feature standard deviation to the standardized posture feature to obtain the posture feature of the first target portrait; The first target portrait is generated according to the expression feature of the first target portrait and the posture feature of the first target portrait.

2. The method according to claim 1, characterized in that The portrait posture control model includes an encoder and a decoder; the acquiring of the expression features of the original first portrait and the posture features of the original first portrait includes: Mapping the original first portrait to a three-dimensional latent space by the encoder, generating three-dimensional key points in the three-dimensional latent space, and encoding and generating expression features and posture features of the original first portrait based on position coordinates of the three-dimensional key points; Generating the target first portrait according to the expression feature of the target first portrait and the posture feature of the target first portrait includes: The expression features of the first target portrait and the posture features of the first target portrait are input into the decoder, and the expression features of the first target portrait and the posture features of the first target portrait are upsampled and decoded by the decoder to output the first target portrait.

3. The method according to claim 1 or 2, characterized in that The portrait posture control model includes an encoder and a decoder; the portrait posture control model is trained by the following steps: Inputting each of the sample video clips in the sample video clip set into a portrait posture control model to be trained, and obtaining expression features and posture features of multiple frames of sample portraits contained in each of the sample video clips through an encoder of the portrait posture control model; Inputting the expression features and posture features of the multiple frames of sample portraits into a decoder of the portrait posture control model, and generating a reconstructed video segment corresponding to each of the sample video segments through the decoder; The portrait pose control model is trained based on the difference between the reconstructed video segment and the sample video segment to obtain the pre-trained portrait pose control model.

4. The method according to claim 3, characterized in that Inputting the expression features and posture features of the multiple frames of sample portraits into a decoder of the portrait posture control model, and generating a reconstructed video segment corresponding to each of the sample video segments through the decoder, comprises: Inputting the expression features and posture features of each frame of the sample portrait contained in the current sample video segment into the decoder, and obtaining the reconstructed portrait of each frame corresponding to the current sample video segment through the decoder; the current sample video segment is any one of the sample video segments in the sample video segment set; The reconstructed portraits of the frames corresponding to the current sample video segment are combined to obtain a reconstructed video segment corresponding to the current sample video segment.

5. A portrait posture control device, characterized in that: The device comprises: an original portrait acquisition module, configured to acquire an original first portrait of a portrait posture to be controlled, and a reference video clip for performing portrait posture control on the original first portrait; the reference video clip carries multiple frames of a second portrait; a posture parameter acquisition module, configured to input the original first portrait and the reference video clip into a pre-trained portrait posture control model, obtain posture features of each second portrait in the reference video clip through the portrait posture control model, and obtain posture control parameters based on the posture features of each second portrait; the posture control parameters including: a posture feature mean and a posture feature standard deviation obtained based on the posture features of each second portrait; An expression parameter acquisition module is configured to acquire expression features of each sample portrait in a sample video clip set, and acquire expression control parameters based on the expression features of each sample portrait; the sample video clip set is a set of sample video clips used to train the portrait posture control model, each sample video clip carrying multiple frames of sample portraits; the expression control parameters include: an expression feature mean and an expression feature standard deviation obtained based on the expression features of each sample portrait; a target portrait generation module, configured to control the original first portrait using the expression control parameters and the posture control parameters to generate a target first portrait after portrait posture control; wherein the expression control parameters are used to control the expression of the target first portrait, and the posture control parameters are used to control the posture of the target first portrait; The target portrait generation module is further configured to obtain expression features of the original first portrait and posture features of the original first portrait; perform standardization processing on the expression features of the original first portrait and the posture features of the original first portrait to obtain standardized expression features and standardized posture features; adjust the standardized expression features using the expression control parameters to obtain expression features of the target first portrait, and adjust the standardized posture features using the posture control parameters to obtain posture features of the target first portrait; and generate the target first portrait based on the expression features of the target first portrait and the posture features of the target first portrait; The target portrait generation module is further used to map the expression feature mean and the expression feature standard deviation to the standardized expression feature to obtain the expression feature of the target first portrait; and map the posture feature mean and the posture feature standard deviation to the standardized posture feature to obtain the posture feature of the target first portrait.

6. The device according to claim 5, characterized in that The portrait posture control model includes an encoder and a decoder; a target portrait generation module is further configured to map the original first portrait to a three-dimensional latent space via the encoder, generate three-dimensional key points in the three-dimensional latent space, and encode and generate expression features and posture features of the original first portrait based on the position coordinates of the three-dimensional key points; The expression features of the first target portrait and the posture features of the first target portrait are input into the decoder, and the expression features of the first target portrait and the posture features of the first target portrait are upsampled and decoded by the decoder to output the first target portrait.

7. The device according to claim 5 or 6, characterized in that The portrait posture control model includes an encoder and a decoder; the device also includes: a control model training module, which is used to input each sample video clip in the sample video clip set into the portrait posture control model to be trained, and obtain the expression features and posture features of multiple frames of sample portraits contained in each sample video clip through the encoder of the portrait posture control model; input the expression features and posture features of the multiple frames of sample portraits into the decoder of the portrait posture control model, and generate reconstructed video clips corresponding to each sample video clip through the decoder; based on the differences between the reconstructed video clips and the sample video clips, train the portrait posture control model to obtain the pre-trained portrait posture control model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Spherical-bottom device and method for tracking video target and controlling posture of video target

    CN102541085A

  • Multi-posture portrait animation generation method and device, equipment and storage medium

    CN119359870A