Image processing method, device and electronic equipment
By acquiring and integrating the appearance and action characteristics of the picture, using optical flow and body attention networks, the problems of unnatural pictures and body loss in complex action scenes in the prior art are solved, and the target picture of natural, real and complete body is achieved.
Patent Information
- Application Number
- CN202011639117.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-12-31
AI Technical Summary
The existing action migration technology is difficult to generate natural and real pictures in complex action scenes, and limb loss is prone to occur.
By obtaining the appearance and action features of the source image, as well as the action features of the target image, calculating displacement information and guidance information, performing feature fusion, and generating target images. This method uses optical flow networks and body attention networks to ensure the integrity and nature of the motion characteristics.
It effectively avoids the problem of limb loss, and the generated target pictures are more natural and realistic, and can maintain complete limb characteristics in complex action scenes.
Smart Images

Figure CN112668517B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motion migration, and in particular to a picture processing method, device and electronic equipment. Background Art
[0002] Motion transfer is an important technology in the field of computer vision and is widely used in many fields such as film production, virtual fitting, and image editing. Given a source image and its corresponding human motion, as well as a target human motion, the motion transfer task is to move the source image to the target motion while maintaining the identity features of the source image. Existing motion transfer tasks are difficult to generate natural and realistic images in some complex motion scenes (such as leg raising, body self-occlusion, hand raising, etc.), and even limb loss may occur. Summary of the invention
[0003] The present invention provides a picture processing method, device and electronic device, so as to solve the problem that the existing action migration technology easily leads to the loss of some features to a certain extent.
[0004] In a first aspect of the present invention, a method for processing an image is provided, the method comprising:
[0005] Acquire appearance features and first action features of the first image, and acquire second action features of the second image;
[0006] According to the appearance feature, the first motion feature and the second motion feature, obtaining displacement information from the first motion feature to the second motion feature and guidance information for guiding feature fusion of the appearance feature and the second motion feature;
[0007] The appearance feature is fused with the second action feature according to the displacement information and the guidance information to generate a target image.
[0008] In a second aspect of the present invention, there is provided a picture processing device, the device comprising:
[0009] A first acquisition module, used to acquire an appearance feature and a first action feature of a first image, and acquire a second action feature of a second image;
[0010] a second acquisition module, configured to acquire, according to the appearance feature, the first action feature, and the second action feature, displacement information from the first action feature to the second action feature, and guidance information for guiding feature fusion of the appearance feature and the second action feature;
[0011] A fusion module is used to fuse the appearance feature with the second action feature according to the displacement information and the guidance information to generate a target image.
[0012] In a third aspect of the present invention, there is also provided an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0013] Memory, used to store computer programs;
[0014] The processor is used to implement the steps in the above-mentioned image processing method when executing the program stored in the memory.
[0015] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the image processing method as described above is implemented.
[0016] In a fifth aspect of the embodiments of the present invention, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute the image processing method as described above.
[0017] Compared with the prior art, the present invention has the following advantages:
[0018] In an embodiment of the present invention, by acquiring the appearance features and the first action features of the first image, and the second action features of the second image, the displacement information from the first action features to the second action features, and the guidance information for guiding the feature fusion of the appearance features and the second action features can be obtained. The first action feature can be transformed into the corresponding position of the second action feature through the displacement information, and the action features lost in the feature fusion process can be supplemented through the guidance information to generate a target image with complete limbs, thereby avoiding the problem of limb missing.
[0019] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for describing the embodiments are briefly introduced below.
[0021] Figure 1 A flow chart of a picture processing method provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of the structure of an action migration model provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of the structure of a limb attention network provided by an embodiment of the present invention;
[0024] Figure 4 A structural block diagram of a picture processing device provided by an embodiment of the present invention;
[0025] Figure 5 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0028] Currently, the existing action migration technologies include the following methods:
[0029] Methods based on decomposition of appearance and action information: This type of method encodes the appearance information and action information separately through the encoder, then connects the two types of information, and generates the target image through the decoder. This type of method does not align the source action with the target action, and the appearance information is difficult to encode into a feature vector, resulting in the generated effect being not realistic enough.
[0030] Methods based on deformation network structures: This type of method uses action structure information to establish a correspondence between the source action and the target action, and deforms the source image through the deformation network structure to generate the target image. However, this type of method has difficulty in handling some complex action scenes.
[0031] Optical flow-based methods: This type of method learns the optical flow information of the source action and the target action through the optical flow network, deforms the source image along the optical flow to the corresponding position of the target action, and combines the target action features to generate a picture of the target action. This method lacks a fine-grained fusion method for the appearance deformation features and the target action features, and is prone to generating unnatural or incomplete limbs in complex action scenes.
[0032] Therefore, the embodiments of the present invention provide an image processing method, device and electronic device, which are based on the motion transfer technology of the limb attention mechanism, and can generate a human body with complete limbs while maintaining naturalness and authenticity even in complex action scenes.
[0033] Specifically, Figure 1 As shown, an embodiment of the present invention provides a method for processing an image, and the method specifically includes:
[0034] Step 101, obtaining appearance features and first action features of a first image, and obtaining second action features of a second image.
[0035] Specifically, this method can be applied to action transfer models, such as Figure 2 As shown, the action migration model includes an appearance feature encoding network 21, which is a downsampling convolution encoding network; the first image I s As the source image, it is input into the appearance feature encoding network 21, and the first image I is convolved through a downsampling convolution process (i.e., a process in which the height and width of the image are reduced and the number of channels is increased). s Encode and output an appearance feature map F r , according to the appearance feature map F r You can get the first picture I s The appearance feature of r It is a feature map with three dimensions: height, width and number of channels. The appearance feature map F r Contains the first picture I s Appearance (such as clothes, skin color, hair, etc.) information.
[0036] And, if Figure 2 As shown, the action migration model may also include an action feature encoding network 22, which is a down-sampling convolution encoding network; an action skeleton point detection network may be used to detect the first image I s The key points of the human body and the second picture P t The key points of the human body, the first picture I s The key points of the human body (such as head, shoulders, neck, etc.) are input into the action feature encoding network 22, and after the downsampling convolution process, the first image I sThe action in the image P is encoded and a first action feature map is output. The first action feature can be obtained according to the first action feature map. The first action feature map contains the basic skeleton information of the first action. t The key points of the human body (such as head, shoulders, neck, etc.) are input into the action feature encoding network 22, and after the downsampling convolution process, the second image P t Encode the action in and output the second action feature map F p , according to the second action feature graph F p The second action feature can be obtained. p Contains the basic skeleton information of the second action.
[0037] It should be noted that the second image can be a single image or a frame image in a video, and is not specifically limited here.
[0038] Step 102: acquiring displacement information from the first motion feature to the second motion feature and guidance information for guiding feature fusion of the appearance feature and the second motion feature according to the appearance feature, the first motion feature and the second motion feature.
[0039] Specifically, Figure 2 As shown, the motion migration model can also include an optical flow network 23, which is an encoding-decoding neural network. It takes the appearance features, the human body key points of the first motion features, and the human body key points of the second motion features as inputs, and through the encoding-decoding process, the displacement information f from the first motion feature to the second motion feature can be obtained.
[0040] Furthermore, the displacement information may be an optical flow vector including three dimensions of height, width and number of channels. For example, the number of channels is 2, i.e., two coordinates of x and y.
[0041] In addition, the action transfer model may also include a limb attention network 24, which is an encoding-decoding neural network. After the encoding-decoding process, a guidance information M can be obtained. The guidance information M is used to guide the feature fusion of the appearance features of the first image and the second action features of the second image, thereby avoiding the problem of loss of action features during the fusion process.
[0042] Furthermore, the guidance information may be a limb attention weight mask including two dimensions of height and width.
[0043] Step 103: According to the displacement information and the guidance information, the appearance feature and the second action feature are fused to generate a target image.
[0044] Specifically, Figure 2 As shown, the action transfer model may also include a decoding network 25, which is an up-sampling convolutional decoding network structure, and the network output is a fused target image I gen The first action feature can be transformed into the corresponding position of the second action feature through the displacement information f, and the action features lost in the feature fusion process can be supplemented through the guidance information M to generate a target image I with a complete limb gen , thereby avoiding problems such as limb loss.
[0045] It should be noted that the application scenarios of the above method are: it can be used in various applications or products that require motion migration, such as virtual fitting, movie making, and dance video generation; specifically, it can have the following functions:
[0046] Function 1: Given a source image (i.e., the first image) and a target action image (i.e., the second image), migrate the source image according to the action of the target action image (i.e., the second action feature) to generate a target image, so that the generated target image has the identity information of the first image (features such as face and clothes), and maintains the action information of the target action (the second action feature).
[0047] Function 2: Given an action sequence and a source picture (i.e., the first picture) of a first video (i.e., the second picture is a frame picture, and multiple frame pictures are combined into the first video), the source picture is replaced with each frame picture of the target video (i.e., the second picture) in the manner of Function 1, thereby generating a target video that retains the identity information of the source picture (i.e., the target picture is a target frame picture, and multiple target frame pictures are combined into the target video).
[0048] In the above embodiment of the present invention, by acquiring the appearance features and the first action features of the first image, and the second action features of the second image, the displacement information from the first action features to the second action features, and the guidance information for guiding the feature fusion of the appearance features and the second action features can be obtained. The first action feature can be transformed into the corresponding position of the second action feature through the displacement information, and the action features lost in the feature fusion process can be supplemented through the guidance information to generate a target image with complete limbs, thereby avoiding the problem of limb missing.
[0049] Optionally, the step 102 of acquiring displacement information from the first action feature to the second action feature according to the appearance feature, the first action feature, and the second action feature may specifically include:
[0050] Encoding the appearance feature, the first motion feature, and the second motion feature to obtain an optical flow feature;
[0051] The optical flow feature is decoded to obtain displacement information from the first motion feature to the second motion feature.
[0052] Specifically, Figure 2 As shown, the appearance feature, the human body key points of the first action feature, and the human body key points of the second action feature are input into the optical flow network 23. After the encoding process, the optical flow feature F corresponding to the optical flow network 23 can be obtained. f ; The optical flow feature F f By decoding, the displacement information f from the first action to the second action, that is, the optical flow vector from the first action feature to the second action feature, can be obtained.
[0053] Optionally, the step 102 of acquiring guidance information for guiding feature fusion of the appearance feature and the second action feature according to the appearance feature, the first action feature, and the second action feature may specifically include:
[0054] Step A1, encoding the first action feature and the second action feature to obtain a limb feature;
[0055] Step A2: obtaining guidance information for guiding feature fusion of the appearance feature and the second action feature according to the optical flow feature and the limb feature.
[0056] Specifically, Figure 2 As shown, the human body key points of the first action feature and the human body key points of the second action feature are input into the limb attention network 24, and the limb feature F is obtained through the encoding process. J , the limb feature F J Contains limb structure information; in the decoding process, through the optical flow feature F f and limb features F J , the guiding information M for guiding the feature fusion of the appearance feature and the second action feature can be obtained, that is, the limb attention weight mask for guiding the feature fusion of the appearance feature and the second action feature can be obtained. The limb attention weight mask is used to indicate whether there is a missing optical flow vector, and locate the missing relationship to each specific position of the second action feature, so as to supplement the feature information and generate the target image I of the limb completion. gen .
[0057] Optionally, the step A2 obtains guidance information for guiding feature fusion of the appearance feature and the second action feature according to the optical flow feature and the limb feature, which may specifically include:
[0058] Step B1, connecting the first channel number of the optical flow feature and the second channel number of the limb feature to obtain a limb weight.
[0059] Specifically, Figure 3 As shown, during the decoding process, the optical flow feature F corresponding to the optical flow network f F J A limb attention network is added between them. Through the limb attention network, the optical flow feature F f The first channel number and the limb feature F J The second channel number is connected to obtain the limb weight a J , the limb weight a J is the limb weight containing two dimensions: height and width.
[0060] For example, the process of obtaining limb weights through the limb attention network is as follows:
[0061] The optical flow feature F f and limb features F J Perform a vector convolution operation conv, then a linear rectified unit (ReLU) operation, then a conv operation, and finally a softmax function operation to obtain the limb weights.
[0062] Step B2: acquiring guidance information for guiding feature fusion of the appearance feature and the second action feature according to the limb weight and the optical flow feature.
[0063] Furthermore, the step B2 described above of obtaining guidance information for guiding the feature fusion of the appearance feature and the second action feature according to the limb weight and the optical flow feature may specifically include:
[0064] Performing a dot product of the limb weight and the optical flow feature to obtain a limb attention feature;
[0065] The third channel number of the limb attention feature is connected to the second channel number of the limb feature, and through an upsampling convolution process, guidance information for guiding feature fusion of the appearance feature and the second action feature is obtained.
[0066] Specifically, Figure 2 and Figure 3 As shown, by changing the limb weight a J and the optical flow feature F f By performing dot multiplication, we can get the limb attention feature The details are as follows:
[0067]
[0068] in, Indicates body attention characteristics;
[0069] F f Represents optical flow features;
[0070] a J represents limb weight;
[0071] Represents dot product.
[0072] And, the body attention feature The third channel number and the limb feature F J The second channel number of the image is connected to perform an upsampling convolution process, and the guiding information M for guiding the feature fusion of the appearance feature and the second action feature can be obtained, that is, the limb attention weight mask for guiding the feature fusion of the appearance feature and the second action feature is obtained, wherein the coordinate value of each position can be a value between 0 and 1, which is used to guide the feature fusion, and the action features lost in the feature fusion process can be supplemented to generate a target image I with a complete limb. gen , thereby avoiding problems such as limb loss.
[0073] Optionally, step 103 may fuse the appearance feature with the second action feature according to the displacement information and the guidance information to generate a target image, which may specifically include:
[0074] Step C1, coordinate deformation of the appearance feature is performed using the displacement information to obtain a deformation feature.
[0075] Specifically, Figure 2 As shown, in the decoding network 25, the appearance feature is first deformed coordinate by coordinate through the displacement information f (i.e., the optical flow vector), thereby obtaining the deformed feature F warp .
[0076] Step C2: Fusing the deformation feature and the second motion feature through the guidance information to obtain a fused feature.
[0077] Specifically, Figure 2 As shown, the deformation feature F warp The second action feature is fused with the guidance information M (i.e., the limb attention weight mask) to obtain the fused feature F fuse .
[0078] Furthermore, the step C2 described above of fusing the deformation feature and the second motion feature through the guidance information to obtain a fused feature may specifically include:
[0079] Performing a dot multiplication process on the guidance information and the deformation feature to obtain a first target feature;
[0080] Performing a dot multiplication process on a second value obtained by subtracting the guidance information from the first value and the second action feature to obtain a second target feature;
[0081] The sum of the first target feature and the second target feature is calculated to obtain a fusion feature.
[0082] The fusion feature can be obtained through the above fusion process. The fusion process adds a limb attention weight mask, and the deformation feature and the second action feature can be fused under the guidance of the limb attention weight mask, which can avoid the loss of features in the fusion process. Among them, the acquisition step of the first target feature and the acquisition step of the second target feature are not limited in sequence. For example: when the first value is 1, the fusion feature can be obtained specifically by the following formula:
[0083]
[0084] Among them, F fuse Indicates fusion features;
[0085] M represents the guidance information, i.e., the limb attention weight mask;
[0086] F warp Indicates deformation characteristics;
[0087] F p Indicates the second action feature;
[0088] Represents dot product.
[0089] Step C3, the fused features are subjected to an upsampling convolution process to generate a target image.
[0090] Specifically, the fused features are subjected to an upsampling convolution process through an upsampling convolutional neural network to be restored into a target image, where the target image has the appearance features of the first image and the second action features of the second image.
[0091] In the training process of the above action transfer model, two pictures of the same person with different actions can be taken (I s , I t ), and the corresponding action features (P s , P t ), the image I after the model generates action transfer gen ; The training process is as follows:
[0092] Step 1: Training the Discriminator
[0093] First, I twith I gen After passing through the discriminator, the adversarial loss is calculated; then the gradient is solved and the weights of the discriminator are updated.
[0094] Step 2: Training the Generator
[0095] First, through I t with I gen , calculate the reconstruction loss and adversarial loss; then I t with I gen Using convolutional neural network (Visual Geometry Group, VGG) to get vgg (I t ) and vgg(I gen ), and calculate the perceptual loss. Then, the target face position is obtained through the face key points, and the face loss (face perceptual loss and face reconstruction loss) is calculated. Then, I s with I t Using the VGG network to get vgg(I s ) and vgg(I t ), use (vgg(I s ), vgg(I t ), f), calculate the optical flow loss. Finally, solve the gradient and update the weight of the generator. Where f is the optical flow vector.
[0096] In the third step, the first and second steps are repeated alternately until the action transfer model converges.
[0097] Specifically, the purpose of the loss function in the above three steps is to provide guidance for the model training process. The training process continuously optimizes the model parameters to make the loss function smaller, thereby learning the ability of action transfer. The data required to complete each training includes two pictures of the same person in different actions (I s , I t ), and the corresponding action features (P s , P t ), the input of the model is (I s , P s , P t ), the output is I gen , I t For I gen The corresponding true value. The above loss function is a joint loss function, and its main components are as follows:
[0098] Reconstruction loss L rec :Let the generated image I gen And the true value image I t Close at the pixel level, expressed as:
[0099] L rec =||Igen -I t ||1
[0100] Among them, the above formula represents the reconstruction loss function L rec For I gen with I t The absolute value of the difference.
[0101] Perceptual loss L perc :Let the generated image I gen and the true value image I t By using a trained VGG neural network to gen and I t Extract features and then calculate the distance between two features, expressed as:
[0102] L per =||vgg(I gen )-vgg(I t )||1
[0103] Among them, vgg(X) represents the X feature extracted by VGG neural network, and X is I gen or I t ;
[0104] Adversarial loss L GAN : Through adversarial loss, the generated images can be made more realistic and natural.
[0105] Optical flow loss L flow :By using the trained VGG neural network to s and the true value image I t Extract features and obtain feature maps vgg(I s ) and vgg(I t ), making vgg(I s )The feature map obtained by pixel-by-pixel deformation along the optical flow vector f is consistent with vgg(I t ), the loss function measures the similarity by cosine distance, which can be expressed as
[0106]
[0107] Among them, φ(*) is the process of deformation of the feature map along the optical flow vector;
[0108] cos(*) is the cosine distance;
[0109] vgg(I t ) l is the feature map vgg(I t ) at the l coordinate position;
[0110] N is the total number of feature map coordinate positions.
[0111] Face loss L face : Determine the face area through the face key points in the target action information, add reconstruction loss and perception loss to the face area separately, and add a separate discriminator for the face area.
[0112] L face =||face(I gen) -face(I t) ||1+||vgg(face(I gen) )-vgg(face(I t) )||1
[0113] Among them, face(*) represents the face area.
[0114] In summary, the joint loss function can be expressed as
[0115] L total =L rec +L prec +L GAN +L flow +L face
[0116] Testing process: If the training in the third step converges, that is, the training is completed, the action transfer model can be used to perform action transfer operations on the first and second input images to generate the target image.
[0117] To summarize, the embodiments of the present invention extract the appearance features of the first image, the first action features of the first image, and the second action features of the second image, obtain the optical flow vector from the first action features to the second action features according to the optical flow network, and obtain the limb attention weight mask through the limb attention network; deform the appearance features along the optical flow vector to obtain the deformation features, and under the guidance of the limb attention weight mask, fuse the deformation features and the second action features, and decode the fused features through the decoding network to generate the target image after the action migration, which can supplement some features lost in the fusion process, ensure the generation of the target image with complete limbs, and make the target image more realistic and natural; and, the addition of face loss in the model training process, for the input low-definition image, a real face image can also be generated through the action migration of the model.
[0118] like Figure 4 As shown, an image processing device 400 provided in an embodiment of the present invention includes:
[0119] A first acquisition module 401 is used to acquire the appearance feature and the first action feature of the first image, and acquire the second action feature of the second image;
[0120] A second acquisition module 402, configured to acquire displacement information from the first action feature to the second action feature and guidance information for guiding feature fusion of the appearance feature and the second action feature according to the appearance feature, the first action feature and the second action feature;
[0121] The fusion module 403 is used to fuse the appearance feature with the second action feature according to the displacement information and the guidance information to generate a target image.
[0122] Optionally, the second acquisition module 402 includes:
[0123] A first encoding unit, configured to obtain an optical flow feature by encoding the appearance feature, the first motion feature, and the second motion feature;
[0124] The first decoding unit is used to obtain displacement information from the first action feature to the second action feature by decoding the optical flow feature.
[0125] Optionally, the second acquisition module 402 further includes:
[0126] A second encoding unit, used for encoding the first action feature and the second action feature to obtain a limb feature;
[0127] An acquisition unit is used to acquire guidance information for guiding feature fusion of the appearance feature and the second action feature according to the optical flow feature and the limb feature.
[0128] Optionally, the acquisition unit includes:
[0129] A connection subunit, used for connecting the first channel number of the optical flow feature and the second channel number of the limb feature to obtain a limb weight;
[0130] An acquisition subunit is used to acquire guidance information for guiding feature fusion of the appearance feature and the second action feature according to the limb weight and the optical flow feature.
[0131] Optionally, the acquisition subunit includes:
[0132] Performing a dot product of the limb weight and the optical flow feature to obtain a limb attention feature;
[0133] The third channel number of the limb attention feature is connected to the second channel number of the limb feature, and through an upsampling convolution process, guidance information for guiding feature fusion of the appearance feature and the second action feature is obtained.
[0134] Optionally, the fusion module 403 includes:
[0135] A deformation unit, used for performing coordinate deformation on the appearance feature through the displacement information to obtain a deformation feature;
[0136] a fusion unit, configured to fuse the deformation feature and the second motion feature through the guidance information to obtain a fused feature;
[0137] The generating unit is used to generate a target image by performing an upsampling convolution process on the fused features.
[0138] Optionally, the fusion unit comprises:
[0139] A first processing subunit, configured to perform a dot multiplication process on the guidance information and the deformation feature to obtain a first target feature;
[0140] A second processing subunit is used to perform a dot multiplication process on the second action feature and a second value obtained by subtracting the guidance information from the first value to obtain a second target feature;
[0141] The calculation subunit is used to calculate the sum of the first target feature and the second target feature to obtain a fusion feature.
[0142] Optionally, the displacement information is an optical flow vector including three dimensions of height, width and number of channels.
[0143] Optionally, the guidance information is a limb attention weight mask including two dimensions of height and width.
[0144] It should be noted that the embodiment of the image processing device is a device corresponding to the above-mentioned image processing method. All implementation methods of the above-mentioned method embodiment are applicable to the embodiment of the device and can achieve the same technical effects, which will not be repeated here.
[0145] To summarize, the embodiments of the present invention extract the appearance features of the first image, the first action features of the first image, and the second action features of the second image, obtain the optical flow vector from the first action features to the second action features according to the optical flow network, and obtain the limb attention weight mask through the limb attention network; deform the appearance features along the optical flow vector to obtain the deformation features, and under the guidance of the limb attention weight mask, fuse the deformation features and the second action features, and decode the fused features through the decoding network to generate the target image after the action migration, which can supplement some features lost in the fusion process, ensure the generation of the target image with complete limbs, and make the target image more realistic and natural; and, the addition of face loss in the model training process, for the input low-definition image, a real face image can also be generated through the action migration of the model.
[0146] The embodiment of the present invention also provides an electronic device. Figure 5 As shown, it includes a processor 501 , a communication interface 502 , a memory 503 and a communication bus 504 , wherein the processor 501 , the communication interface 502 , and the memory 503 communicate with each other via the communication bus 504 .
[0147] The memory 503 is used to store computer programs.
[0148] When the processor 501 is used to execute the program stored in the memory 503, part or all of the steps in the image processing method provided by the embodiment of the present invention are implemented.
[0149] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0150] The communication interface is used for communication between the above terminal and other devices.
[0151] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0152] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0153] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer executes the image processing method described in the above embodiment.
[0154] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes the image processing method described in the above embodiment.
[0155] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0156] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for processing an image, characterized in that: The method comprises: Acquire appearance features and first action features of the first image, and acquire second action features of the second image; According to the appearance feature, the first motion feature and the second motion feature, obtaining displacement information from the first motion feature to the second motion feature and guidance information for guiding feature fusion of the appearance feature and the second motion feature; According to the displacement information and the guidance information, the appearance feature and the second action feature are fused to generate a target image; The acquiring, according to the appearance feature, the first action feature and the second action feature, guidance information for guiding the feature fusion of the appearance feature and the second action feature includes: Encoding the appearance feature, the first motion feature, and the second motion feature to obtain an optical flow feature; The first action feature and the second action feature are encoded to obtain a limb feature; According to the optical flow feature and the limb feature, obtaining guidance information for guiding feature fusion of the appearance feature and the second action feature through a limb attention network; The step of obtaining, according to the optical flow feature and the limb feature, guidance information for guiding the feature fusion of the appearance feature and the second action feature through a limb attention network includes: Connecting the first channel number of the optical flow feature and the second channel number of the limb feature to obtain a limb weight; Performing a dot product of the limb weight and the optical flow feature to obtain a limb attention feature; The third channel number of the limb attention feature is connected to the second channel number of the limb feature, and through an upsampling convolution process, guidance information for guiding feature fusion of the appearance feature and the second action feature is obtained.
2. The method according to claim 1, characterized in that Acquiring displacement information from the first action feature to the second action feature according to the appearance feature, the first action feature, and the second action feature, including: The optical flow feature is decoded to obtain displacement information from the first motion feature to the second motion feature.
3. The method according to claim 1, characterized in that The step of fusing the appearance feature with the second action feature according to the displacement information and the guidance information to generate a target image includes: Performing coordinate deformation on the appearance feature by using the displacement information to obtain a deformation feature; Fusing the deformation feature and the second motion feature using the guidance information to obtain a fused feature; The fused features are passed through an upsampling convolution process to generate a target image.
4. The method according to claim 3, characterized in that The step of fusing the deformation feature and the second motion feature through the guidance information to obtain a fused feature includes: Performing a dot multiplication process on the guidance information and the deformation feature to obtain a first target feature; Performing a dot multiplication process on a second value obtained by subtracting the guidance information from the first value and the second action feature to obtain a second target feature; The sum of the first target feature and the second target feature is calculated to obtain a fusion feature.
5. The method according to claim 1, characterized in that: The displacement information is an optical flow vector including three dimensions: height, width and number of channels.
6. The method according to claim 1, characterized in that The guidance information is a limb attention weight mask including two dimensions of height and width.
7. A picture processing device, characterized in that: The device comprises: A first acquisition module, used to acquire an appearance feature and a first action feature of a first image, and acquire a second action feature of a second image; a second acquisition module, configured to acquire, according to the appearance feature, the first action feature, and the second action feature, displacement information from the first action feature to the second action feature, and guidance information for guiding feature fusion of the appearance feature and the second action feature; The second acquisition module includes: A first encoding unit, configured to obtain an optical flow feature by encoding the appearance feature, the first motion feature, and the second motion feature; A second encoding unit, used for encoding the first action feature and the second action feature to obtain a limb feature; An acquisition unit, configured to acquire, according to the optical flow feature and the limb feature, guidance information for guiding feature fusion of the appearance feature and the second action feature; a fusion module, configured to fuse the appearance feature with the second action feature according to the displacement information and the guidance information to generate a target image; The acquisition unit includes a submodule for connecting the first channel number of the optical flow feature and the second channel number of the limb feature to obtain a limb weight, performing a dot product between the limb weight and the optical flow feature to obtain a limb attention feature, connecting the third channel number of the limb attention feature and the second channel number of the limb feature, and obtaining guidance information for guiding the feature fusion of the appearance feature and the second action feature through an upsampling convolution process.
8. An electronic device, characterized in that: include: A processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the steps in the image processing method as described in any one of claims 1 to 6 when executing the program stored in the memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the image processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Human body motion animation generating method and system of two-dimensional virtual figure
CN108230431A
Image editing method and device, electronic equipment and storage medium
CN111814566A