A Deep Learning System and Method for Video Action Transfer
By extracting and quantizing the key points and depth features of video images, combined with the video action transfer deep learning system optimized by loss function, the problem of unreality and fuzzy action transfer in the existing technology is solved, and the high-definition action transfer effect is achieved.
Patent Information
- Application Number
- CN202111660869.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-30
AI Technical Summary
In the video action migration, it is difficult to achieve large-scale action migration without relying on complex human models, and the existing methods are prone to problems such as unrealistic actions and blurred action areas.
By extracting key point information and depth features of source and reference images, quantizing the process, using video action transfer deep learning systems and methods, reconstructing the target image, combining global and local quantization features, using loss functions to optimize the feature recombination network, and generating target images with high clarity.
The action migration process is simplified, the clarity and resolution of the target image is improved, and the authenticity and effect of action migration can be effectively completed especially when the posture changes greatly.
Smart Images

Figure CN114399708B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a deep learning system and method for video action migration. Background Art
[0002] With the rise of the video era, portrait video generation has also become an important task in the field of computer vision, and has extensive application scenarios in fields such as video directed editing and video production. Among them, the action migration task in the video aims to transfer the actions of the reference video character to the source video character while retaining the identity characteristics of the source video character. This task has received great attention in the fields of animation and film, and has great practical application value.
[0003] Currently, some classic methods rely on complex human body modeling processes, such as human contour models, face models (Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pages 187 - 194, 1999.), etc., to complete the generation of action migration videos on these highly restricted and complex models. When the application scenario changes, it is difficult for these classic methods to be migrated to new datasets. With the rise of deep models such as autoencoders and generative adversarial networks (David Berthelot, Thomas Schumm, and Luke Metz. Be - gan: Boundary equilibrium generative adversarial networks. arXiv preprint arXiv:1703.10717, 2017.), these models have also been used for video generation tasks, but these methods cannot flexibly control the actions and appearances of the generated video characters.
[0004] On the other hand, there are also methods (Siarohin A, Lathuilière S, Tulyakov S, et al. Animating arbitrary objects via deep motion transfer [C] Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 2377-2386.) that use predicted motion optical flow and warp the source video characters to achieve motion transfer. However, in the case of large motion transformations, such methods are prone to generating motion transfer videos with unrealistic motions, blurred motion regions, and unnatural visuals.
[0005] Therefore, how to achieve large-scale motion transfer without the need for complex human models is an urgent problem to be solved at present. Summary of the Invention
[0006] The object of the present invention is to provide a deep learning system and method for video motion transfer, which extracts key point information of the source image, key point information of the reference image, deep features of the source image, and deep features of the reference image, quantifies the deep features of the source image and the deep features of the reference image respectively to obtain the global quantization features of the source image and the global quantization features of the reference image. Based on the global quantization features of the source image, the deep features of the reference image are quantified again to obtain the local quantization features of the reference image. Based on the global quantization features of the reference image, the deep features of the source image are quantified again to obtain the local quantization features of the source image. The key point coordinates of the source image character, the global quantization features of the source image, and the local quantization features of the reference image are predicted to obtain the quantization features of the target image, and the quantization features of the target image are decoded to obtain the target image. The image is reconstructed according to the human key point information and quantization features to ensure the clarity and resolution of the target image.
[0007] In the first aspect, the above object of the present invention is achieved by the following technical solutions:
[0008] A deep learning system for video action transfer, comprising a person video data preprocessing unit, a video feature quantization unit, a video feature recombination unit, and an action transfer video generation unit, all of which are respectively connected to a system control unit. The person video data preprocessing unit is used to preprocess source image data and reference image data, and extract source key point information in the source image and reference key point information in the reference image. The video feature quantization unit is used to extract the depth features of the source image and the reference image respectively, and perform feature quantization operations to obtain the source image quantization features and the reference image quantization features. The video feature recombination unit is used to predict the quantization features of the target image according to the source key points, the source image quantization features, and the reference image quantization features. The action transfer video generation unit is used to output the target image according to the quantization features of the target image. The system control unit is used to store programs and perform control.
[0009] The present invention is further configured as follows: It further includes an input control unit, a video display unit, and a system communication unit, all of which are respectively connected to the system control unit. The system communication unit is used for data interaction between different structural units. The input control unit is used to provide image data input. The video display unit is used to output the action video of the target image.
[0010] In a second aspect, the above object of the present invention is achieved by the following technical solutions:
[0011] A deep learning method for video action transfer, which establishes a video action transfer model, extracts two different frame images from the same video as the source image and the reference image, performs preprocessing, and respectively extracts the source key point information of the source image, the reference key point information of the reference image, the source image depth feature, and the reference image depth feature. The depth features of the source image and the reference image are respectively quantified to obtain the source image quantization features and the reference image quantization features. Prediction is performed according to the source key point information, the reference key point information, the source image quantization features, and the reference image quantization features to obtain the target image quantization features. The target image is generated according to the target image quantization features. Based on the parameters of the video action transfer model, for the transfer source images and transfer reference images from different sources, the same process as that in modeling is adopted for action transfer.
[0012] The present invention is further configured as follows: The depth feature of the source image is quantified to obtain the global quantization feature of the source image, and the depth feature of the reference image is quantified to obtain the global quantization feature of the reference image. Based on the global quantization feature of the reference image, the depth feature of the source image is quantified again to obtain the local quantization feature of the source image. Based on the global quantization feature of the source image, the depth feature of the reference image is quantified again to obtain the local quantization feature of the reference image.
[0013] The present invention is further configured to: calculate the minimum Euclidean distance of each feature in the depth features of the source image in the global quantization features of the reference image to obtain the local quantization features of the source image; calculate the minimum Euclidean distance of each feature in the depth features of the reference image in the global quantization features of the source image to obtain the local quantization features of the reference image.
[0014] The present invention is further configured to: perform prediction according to the source key point information, the reference key point information, the global quantization features of the source image, and the local quantization features of the reference image to obtain the quantization features of the target image.
[0015] The present invention is further configured to: map the source key point information and the reference key point information into dimensional features to obtain a key point feature sequence, respectively establish index sequences according to the global quantization features of the source image and the local quantization features of the reference image, and based on the key point feature sequence and the index sequences, in the feature recombination network, adopt a masking method to obtain the highest probability value after masking, establish the index sequence of the target image, and predict the quantization feature index of the target image in the source image index sequence.
[0016] The present invention is further configured to: calculate the probability distribution value of the quantization feature index of the target image and the index loss of the local quantization features of the reference image to obtain a second loss function, and optimize the feature recombination network.
[0017] The present invention is further configured to: establish a global feature library, an encoder, and a decoder, use the encoder to extract the depth features of the source image and the depth features of the reference image respectively, use the global feature library to quantize the depth features of the source image and the depth features of the reference image respectively, and use the decoder to decode the quantization features of the target image to generate the target image.
[0018] The present invention is further configured to: set and initialize the global feature library according to the image data, and optimize the global feature library with the minimum Euclidean distance between the depth features and the quantization features; use PatchGAN to distinguish between the generated image and the real image, and use the first loss function to train the encoder, decoder, and global feature library at the same time. The first loss function includes adversarial loss and quantization recombination loss.
[0019] The encoder includes a convolutional layer, a residual module, a downsampling module, a self-attention module, and an activation function, and is used to integrate and transform the pixel features of the original image and perform mapping to obtain an intermediate feature map. The decoder is symmetric to the encoder and includes a convolutional layer, a residual module, an upsampling module, a self-attention module, and an activation function, and is used to decode the quantization features of the target image.
[0020] The present invention is further configured to: preprocess the source image and the reference image, including resizing the images, using a pre-trained model and a data augmentation method, and obtaining the coordinates of the person key points of the source image and the coordinates of the person key points of the reference image.
[0021] In a third aspect, the above-mentioned object of the present invention is achieved by the following technical solutions:
[0022] A deep learning terminal for video action migration includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in this application is implemented.
[0023] In a fourth aspect, the above-mentioned object of the present invention is achieved by the following technical solutions:
[0024] A computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method described in this application is implemented.
[0025] Compared with the prior art, the beneficial technical effects of this application are as follows:
[0026] 1. In this application, by quantifying the deep features of the source image and the reference image, and combining the key point information to predict the quantified features of the target image, and generating the target image after decoding, the process of action migration is simplified, and the clarity and resolution of the target image are improved.
[0027] 2. Further, by introducing a loss function in this application, the data processing is more realistic, and the effect of action migration is improved.
[0028] 3. Further, by extracting the human key point information as the condition of the feature recombination network in this application, the action migration with large pose changes can be effectively completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic structural diagram of a migration deep learning system according to a specific embodiment of this application;
[0030] Figure 2 is a schematic structural diagram of a migration deep learning modeling process according to a specific embodiment of this application;
[0031] Figure 3 is a schematic structural diagram of a feature quantization process according to a specific embodiment of this application;
[0032] Figure 4 is a schematic structural diagram of a feature recombination process according to a specific embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The present invention will be further described in detail below with reference to the accompanying drawings. Specific Embodiment 1
[0035] A deep learning system for video action migration in this application is as follows Figure 1 As shown, it includes a preprocessing unit for human video data, a quantization unit for video features, a recombination unit for video features, a generation unit for action migration video, a system control unit, an input control unit, a video display unit, and a system communication unit. The preprocessing unit for human video data, the quantization unit for video features, the recombination unit for video features, the generation unit for action migration video, the input control unit, the video display unit, and the system communication unit are respectively connected to the system control unit.
[0036] The preprocessing unit for human video data is used to preprocess the source image data and the reference image data, and extract the source key point information in the source image and the reference key point information in the reference image, including the source person key point coordinate information in the source image and the reference person key point coordinate information in the reference image.
[0037] The quantization unit for video features is used to extract the depth features of the source image and the reference image respectively, and perform feature quantization operations respectively to obtain the source image quantization features and the reference image quantization features. The source image quantization features include the source image global quantization features and the source image local quantization features, and the reference image quantization features include the reference image global quantization features and the reference image local quantization features.
[0038] The recombination unit for video features is used to predict the quantization features of the target image according to the source key points, the source image quantization features, and the reference image quantization features; it includes recombining with the source person key point coordinates, the reference person key point coordinates, the source image global quantization features, and the reference image local quantization features to obtain the action migration target image quantization features of the source person. Or recombining with the source person key point coordinates, the reference person key point coordinates, the reference image global quantization features, and the source image local quantization features to obtain the action migration target image quantization features of the reference person.
[0039] The generation unit for action migration video is used to output the target image according to the quantization features of the target image, and the system control unit is used to store the program and perform control.
[0040] Extract the image depth features and key point information of all images in the source video and all images in the reference video, and perform feature quantization and recombination to obtain the video quantization features of the target image and generate the target video.
[0041] The system communication unit is used for data interaction between different structural units, the input control unit is used to provide video and image data input, and the video display unit is used to output the action video of the target image or the target video. Specific Embodiment 2
[0043] Video is composed of frames of images. For simplicity, this application uses images to illustrate action migration. For action migration of video, it can be deduced by analogy and will not be elaborated here.
[0044] A deep learning method for video action migration in this application, as Figure 2 shown, a video action migration model is established, and the parameters in each step are updated, including the following steps:
[0045] S1. Start;
[0046] S2. Receive video data, intercept different frame images from the video data, and use them as the source image and the reference image respectively;
[0047] S3. Preprocess the source image and the reference image to obtain the source key point information in the source image and the reference key point information in the reference image;
[0048] S4. Extract the depth features of the source image and the reference image respectively, and quantize the depth features to obtain the global quantization feature and the local quantization feature;
[0049] S5. According to the key point information, the global quantization feature, and the local quantization feature, perform feature recombination to predict the quantization feature of the target image;
[0050] S6. Generate the target image according to the target image quantization feature;
[0051] S7. Display the target image;
[0052] S8. End.
[0053] Before extracting the depth features and quantizing the features, a global feature library, an encoder, and a decoder are established. The global feature library is used to quantize the depth features, the encoder is used to extract the depth features, and the decoder is used to decode the quantized features to generate the target image.
[0054] Set the global feature library according to the size and complexity of the image data, and update the parameters of the encoder, decoder, and global feature library respectively based on the first loss function.
[0055] The first loss function includes the adversarial loss and the quantization recombination loss
[0056]
[0057]
[0058]
[0059] In the formula, x s represents the source image, represents the reconstructed source image after being reconstructed by the encoder and the decoder, x tDenotes the reference image, Denotes the reconstructed reference image after being reconstructed by the encoder and decoder, z s Denotes the depth feature of the source image, Denotes the local quantization feature of the source image, z t Denotes the depth feature of the reference image, Denotes the local quantization feature of the reference image, Is the loss function for measuring the difference between the features in the feature storage module and the depth features of the source image and the reference image.
[0060] β represents the balance coefficient, and its value range is: [0, 1].
[0061] sg represents the gradient stop operation, which is used to update the encoder parameters and the global feature library.
[0062] Utilize the adversarial loss to improve the authenticity of the image reconstructed by the encoder and decoder, and use PatchGAN to distinguish between the generated image and the real image. The quantization recombination loss is used to measure the quality of the image after the quantization feature arrangement is recombined, and the quantization recombination loss is used to update the features in the global feature library.
[0063] The encoder includes a convolutional layer, a residual module, a downsampling module, a self-attention module, and an activation function. The decoder is symmetrically set with the encoder and includes a convolutional layer, a residual module, an upsampling module, a self-attention module, and an activation function.
[0064] When training the encoder and decoder, adopt the first loss function and introduce PatchGAN to improve the authenticity of the reconstructed image.
[0065] For the source image and the reference image obtained from the same video, adjust the size, and use the pre-trained model to extract the key point information respectively. Extract the source key point information from the source image and the reference key point information from the reference image. Introduce the data augmentation method of random flipping of the image to improve the training effect of the model.
[0066] Using the encoder, respectively obtain the depth feature z of the source image S s and the depth feature z of the reference image T t . The depth feature z of the source image S s and the depth feature z of the reference image T t Both have the dimension of H×W×C.
[0067] As Figure 3 shown, search for the library depth feature closest to the depth feature in the global feature library as the global quantization feature. The dimension of the global quantization feature is still H×W×C. The global quantization feature of the source image is marked as Z s , and the global quantization feature of the reference image is marked as Zt 。
[0068] From the global quantization features, local quantization is further performed to obtain local quantization features. Specifically, from the global quantization features Z of the source image s find the quantization feature closest to the depth feature z of the reference image t as the local quantization feature of the reference image Correspondingly, from the global quantization features Z of the reference image t find the quantization feature closest to the depth feature z of the source image s as the local quantization feature of the source image
[0069] Find the H×W library depth features with dimension C closest to the depth feature from the global feature library as the global quantization features, which are respectively expressed as follows:
[0070]
[0071]
[0072] In the formula, q k represents the k-th library depth feature in the global feature library q, and (i, j) respectively represent the coordinates on the depth feature map; represents the depth feature corresponding to the i-th row and j-th column on the depth feature map of the reference image.
[0073]
[0074]
[0075] In the formula, z k represents the k-th quantization feature in the global quantization features; (i, j) respectively represent the coordinates on the depth feature map; represents the depth feature corresponding to the i-th row and j-th column on the depth feature map of the source image.
[0076] Build a feature recombination network, and perform recombination according to the source key point information, reference key point information, global quantization features, and local quantization features to predict the quantization features of the target image as shown in Figure 4 shown.
[0077] Fix the parameters of the encoder and decoder, and use a 3-layer fully connected network with a ReLU activation function to map the key point coordinates in the key point information to the dimension features to obtain the key point feature sequence
[0078] Let the global quantization features Z of the source image s ∈RH×W×C , convert to the global quantization feature index sequence of the source image Local quantization features of the reference image , convert to the local quantization feature index sequence of the reference image
[0079] Using the key-point feature sequence as conditional information, and according to the global quantization feature index sequence of the source image the local quantization feature index sequence of the reference image predict the quantization feature index sequence of the target image
[0080]
[0081] where t [s] is a learnable start index, s is all the indexes of the global quantization feature index sequence of the source image, c is all the features of the key-point feature sequence, t <j is the first j - 1 indexes of the local quantization feature index sequence of the reference image. is the j-th index of the quantization feature index sequence of the target image. When generating the quantization feature index of the target image , first screen out the quantization feature indexes of the source image from the global feature library From predict the quantization feature index of the target image within the range
[0082] In a specific embodiment of the present application, at the Softmax layer at the end of the feature recombination network, in a masked manner, cover the index prediction probability values that do not belong to the range, and obtain the index sequence of the target image according to the highest probability value after masking. By using the masked Softmax layer, complete the recombination of the source image quantization features, and generate the quantization features of the target image from the global quantization features of the source image.
[0083] Calculate the second loss function using the quantization feature index probability distribution value of the target image and the local quantization feature index label of the reference image, and update the feature recombination network.
[0084] The second loss function is:
[0085]
[0086] where s is all the indexes of the global quantization feature index sequence of the source image, c is all the features of the key-point feature sequence, is all the indexes of the quantization feature index sequence of the target image.
[0087] Decode the predicted target image quantization features through a decoder to obtain the target image
[0088] After the above steps, the establishment of the video action migration model is completed, and the parameters in the video action migration model are determined.
[0089] In a specific embodiment of the present application, the encoder includes 2 convolutional layers, 7 residual modules, 4 downsampling modules, 2 self-attention modules, and a Swish activation function. The decoder includes 2 convolutional layers, 7 residual modules, 4 upsampling modules, 2 self-attention modules, and a Swish activation function.
[0090] Search for the closest depth feature from the global feature library, which is to calculate the Euclidean distance between the depth feature in the library and the selected depth feature, and the one with the smallest Euclidean distance is the closest depth feature. Search for the closest depth feature from the global quantization features, which is also to calculate the Euclidean distance between the depth feature in the global quantization features and the selected depth feature, and the one with the smallest Euclidean distance is the closest depth feature.
[0091] The feature recombination network includes a multi-layer Transformer network with the architecture of GPT2. Set the feature embedding dimension of each layer. In the multi-head attention mechanism of each layer, set the number of attention heads. For example, set a 12-layer Transformer network, set the feature embedding dimension to 768, and the number of attention heads to 12.
[0092] Based on the video action migration model parameters, preprocess the migration source images and migration reference images from different sources according to the steps of establishing the video action migration model to obtain the key point information of the migration source images and the key point information of the migration reference images respectively. Based on the encoder, obtain the depth features of the migration source images and the depth features of the migration reference images. Based on the global feature library, quantize the depth features respectively to obtain the global quantization features of the migration source images and the global quantization features of the migration reference images. Based on the global quantization features of the migration source images, re-quantize the depth features of the migration reference images to obtain the local quantization features of the migration reference images. Input the key point information of the migration source images, the key point information of the migration reference images, the global quantization features of the migration source images, and the local quantization features of the migration reference images into the feature recombination network for recombination to obtain the quantization features of the migration target images, and use the decoder to decode to obtain the migration target images.
[0093] Repeat the above operations for each frame image in a migration source video and a migration reference video to obtain a migration target video. Specific Embodiment Three
[0095] A terminal device of a video action migration deep learning system provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an image preprocessing model calculation program. When the processor executes the computer program, the methods in Embodiments 1 and 2 are implemented.
[0096] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device of the video action migration deep learning system. For example, the computer program can be divided into multiple modules, and the specific functions of each module are as follows:
[0097] 1. An image preprocessing module, used for preprocessing images;
[0098] 2. A quantization module, used for quantizing the depth features of images;
[0099] 3. A recombination module, used for generating the quantized features of the target image;
[0100] 4. A generation module, used for generating the target image.
[0101] The terminal device of the video action migration deep learning system can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device of the video action migration deep learning system may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above examples are only examples of the terminal device of the video action migration deep learning system, and do not constitute a limitation on the terminal device of the video action migration deep learning system. It may include more or fewer components, or combine certain components, or different components. For example, the terminal device of the video action migration deep learning system may also include a network access device, a bus, etc.
[0102] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the video action migration deep learning system terminal device, and connects various parts of the entire video action migration deep learning system terminal device through various interfaces and lines.
[0103] The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the video action migration deep learning system terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Specific Embodiment Four
[0105] The module / unit integrated in the terminal device of the video action migration deep learning system, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0106] The embodiments of this specific implementation manner are all preferred embodiments of the present invention, and do not limit the protection scope of the present invention accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A deep learning method for video action transfer, characterized in that: A video action transfer model is established, and two different frame images are extracted from the same video as the source image and the reference image, and preprocessing is performed; The source key point information of the source image, the reference key point information of the reference image, the source image depth feature, and the reference image depth feature are respectively extracted; The source image depth feature and the reference image depth feature are respectively quantified to obtain a source image quantization feature and a reference image quantization feature; Predictions are made according to the source key point information, the reference key point information, the source image quantization feature, and the reference image quantization feature to obtain a target image quantization feature; A target image is generated according to the target image quantization feature; Based on the parameters of the video action transfer model, for transfer source images and transfer reference images from different sources, an action transfer is performed using the same process as modeling.
2. The deep learning method for video action transfer according to claim 1, characterized in that: The source image depth feature is quantified to obtain a source image global quantization feature; The reference image depth feature is quantified to obtain a reference image global quantization feature; Based on the reference image global quantization feature, the source image depth feature is quantified again to obtain a source image local quantization feature; Based on the source image global quantization feature, the reference image depth feature is quantified again to obtain a reference image local quantization feature.
3. The deep learning method for video action transfer according to claim 2, characterized in that: The minimum Euclidean distance of each feature in the source image depth feature in the reference image global quantization feature is calculated to obtain a source image local quantization feature; The minimum Euclidean distance of each feature in the reference image depth feature in the source image global quantization feature is calculated to obtain the reference image local quantization feature.
4. The deep learning method for video action transfer according to claim 3, characterized in that: Predictions are made according to the source key point information, the reference key point information, the source image global quantization feature, and the reference image local quantization feature to obtain a target image quantization feature.
5. The deep learning method for video action transfer according to claim 4, characterized in that: The source key point information and the reference key point information are mapped into dimensional features to obtain a key point feature sequence; Index sequences are respectively established according to the source image global quantization feature and the reference image local quantization feature; Based on the key point feature sequence and the index sequence, in the feature recombination network, a masked highest probability value is obtained by using a masking method; An index sequence of the target image is established, and a target image quantization feature index is predicted in the index sequence of the source image.
6. The deep learning method for video action transfer according to claim 5, characterized in that: The probability distribution value of the target image quantization feature index and the index loss of the reference image local quantization feature are calculated to obtain a second loss function, and the feature recombination network is optimized.
7. The video action transfer deep learning method according to claim 6, wherein: Before extracting the depth features and quantifying the features, a global feature library, an encoder, and a decoder are established. And the encoder is used to extract the source image depth feature and the reference image depth feature respectively; The global feature library is used to quantize the source image depth feature and the reference image depth feature respectively; The decoder is used to decode the target image quantization feature to generate the target image.
8. The video action transfer deep learning method according to claim 7, wherein: The global feature library is set and initialized according to the image data, and the global feature library is optimized with the minimum Euclidean distance between the depth feature and the quantization feature; PatchGAN is used to distinguish the generated image from the real image; The first loss function is used to train the encoder, the decoder and the global feature library at the same time, and the first loss function includes adversarial loss and quantization recombination loss.
9. The video action transfer deep learning method according to claim 8, wherein: The encoder includes a convolutional layer, a residual module, a downsampling module, a self-attention module, and an activation function, and is used to integrate and transform the pixel features of the original image and perform mapping to obtain an intermediate feature map; The decoder is symmetric with the encoder and includes a convolutional layer, a residual module, an upsampling module, a self-attention module, and an activation function, and is used to decode the target image quantization feature.
10. The video action transfer deep learning method according to claim 9, wherein: Preprocessing the source image and the reference image, including resizing the image, using a pre-trained model and a data augmentation method, and obtaining the source image human key point coordinates and the reference image human key point coordinates.
11. A deep learning terminal for video action migration, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1-10 is implemented.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1-10 is implemented.