Image Processing Method and Apparatus
By detecting the posture of the target image object and generating a synthetic image, the problems of poor user experience and low image processing efficiency in the prior art are solved, and the effect of generating a dance video without uploading the original video is achieved.
Patent Information
- Application Number
- CN202110915053.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-08-10
AI Technical Summary
The prior art requires uploading the original video of their own dancing for generator training when users imitate dance videos, resulting in poor user experience, high learning cost and low image processing efficiency.
By acquiring the target image and the action reference video, detecting whether the target image contains the target image and determining whether its posture meets the preset conditions, generating a synthetic image of the reference image object action in the target image object simulation action reference video, and finally performing video synthesis processing to generate the target action video.
Dance videos can be generated without users uploading their own dancing videos, which improves image processing efficiency, enhances user experience, and is highly applicable.
Smart Images

Figure CN113516762B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an image processing method and apparatus. Background Art
[0002] With the continuous development of network applications, various dance games have gradually become rich. Generally speaking, when a user wants to imitate the dance movements of a person in a dance video to generate a video of the user dancing this dance, the user usually needs to first upload an original video of their own dancing, train a generator based on this original video, and then the user inputs their own image and the dance video they want to imitate into this generator to generate a video of the user imitating the dance video to dance. However, this method requires the user to upload an original video of their own dancing to train the generator, resulting in a poor user experience. In addition, for different users, they need to learn the corresponding generator, which has a high learning cost and low image processing efficiency. Summary of the Invention
[0003] Embodiments of this application provide an image processing method and apparatus, which can improve the image processing efficiency, enhance the user experience, and have high applicability.
[0004] In a first aspect, embodiments of this application provide an image processing method, which includes:
[0005] Obtain a target image and determine an action reference video, where the action reference video includes multiple reference images, and each reference image in the multiple reference images includes a reference image object;
[0006] When it is detected that the target image object is included in the target image and it is determined that the posture of the target image object meets a preset condition, obtain a first 3D model image corresponding to the target image object;
[0007] Obtain a second 3D model image corresponding to the reference image object in the target reference image of the action reference video, where the target reference image is one of the multiple reference images;
[0008] Generate a composite image in which the target image object simulates the action of the reference image object in the target reference image according to the target image object, the first 3D model image, and the second 3D model image;
[0009] Perform video composition processing on multiple composite images to obtain a target action video.
[0010] In combination with the first aspect, in a possible implementation manner, the method further includes:
[0011] Perform pose recognition processing on the above target image to obtain key point information of multiple human key points included in the above target image, where the key point information includes position coordinate information and confidence information of the human key points;
[0012] Determine whether the target image includes a target image object based on the above multiple human key point information, and determine whether the pose of the target image object satisfies a preset condition based on the above multiple human key point information.
[0013] Combined with the first aspect, in a possible implementation manner, the above determining whether the target image includes a target image object based on the above multiple human key point information includes:
[0014] Determine the number of human key points in the above multiple human key point information whose confidence information is greater than or equal to a first preset threshold;
[0015] If the number of the above human key points is greater than or equal to a second preset threshold, determine that the target image includes a target image object.
[0016] Combined with the first aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above multiple human key points include a left shoulder, a left hand, a right shoulder, and a right hand;
[0017] The above determining whether the pose of the target image object satisfies a preset condition based on the above multiple human key point information includes:
[0018] If the abscissa value of the above left hand is less than the abscissa value of the above left shoulder, the abscissa value of the above right hand is greater than the abscissa value of the above right shoulder, and the absolute value of the difference between the abscissa value of the above left hand and the abscissa value of the above right hand is greater than a third preset threshold, determine that the pose of the target image object satisfies the preset condition;
[0019] Wherein, the above third preset threshold is determined by the absolute value of the difference between the abscissa value of the above left shoulder and the abscissa value of the above right shoulder.
[0020] Combined with the first aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above multiple human key points include a left foot and a right foot;
[0021] The above determining whether the pose of the target image object satisfies a preset condition based on the above multiple human key point information includes:
[0022] If the abscissa value of the above left foot is less than the abscissa value of the above right foot, determine that the pose of the target image object satisfies the preset condition.
[0023] In combination with the first aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the multiple human key points include the left eye and the right eye;
[0024] Determining whether the pose of the target image object meets a preset condition according to the multiple human key point information includes:
[0025] Determining an angle between a line connecting the left eye and the right eye and the horizontal direction according to the position coordinate information of the left eye and the position coordinate information of the right eye;
[0026] If the angle is less than a fourth preset threshold, it is determined that the pose of the target image object meets the preset condition.
[0027] In combination with the first aspect, in a possible implementation manner, generating a composite image in which the target image object simulates the actions of the reference image object in the target reference image according to the target image object, the first 3D model diagram, and the second 3D model diagram includes:
[0028] Determining a transformation matrix according to the first 3D model diagram and the second 3D model diagram;
[0029] Performing feature extraction processing on the target image object and the first 3D model diagram through a person restoration adversarial neural network, and obtaining first image feature parameters output by each of the n first sampling convolutional networks included in the person restoration adversarial neural network, where n is an integer greater than 1;
[0030] Performing processing on the transformation matrix, the second 3D model diagram, and the first image feature parameters output by each of the first sampling convolutional networks through a picture synthesis adversarial neural network to obtain the composite image.
[0031] In combination with the first aspect, in a possible implementation manner, the picture synthesis adversarial neural network includes n second sampling convolutional networks, and the n second sampling convolutional networks correspond one-to-one to the n first sampling convolutional networks;
[0032] Performing processing on the transformation matrix, the second 3D model diagram, and the first image feature parameters output by each of the first sampling convolutional networks through the picture synthesis adversarial neural network to obtain the composite image includes:
[0033] Inputting the transformation matrix and the second 3D model diagram into the first second sampling convolutional network of the n second sampling convolutional networks to obtain second image feature parameters output by the first second sampling convolutional network;
[0034] Input the second image feature parameters output by the second sampling convolutional network of the i-th layer and the first image feature parameters output by the first sampling convolutional network of the (i + 1)-th layer into the second sampling convolutional network of the (i + 1)-th layer for processing, to obtain the second image feature parameters output by the second sampling convolutional network of the (i + 1)-th layer, until the second image feature parameters output by the second sampling convolutional network of the n-th layer are obtained as the synthesized image, where 1 ≤ i ≤ n - 1 and i is an integer.
[0035] Combined with the first aspect, in a possible implementation, the above video synthesis processing of multiple synthesized images to obtain a target action video includes:
[0036] Obtain the frame rate and audio information of the above action reference video;
[0037] Obtain a background picture, and fuse the above multiple synthesized images with the above background picture to generate multiple fused images;
[0038] Synthesize the above multiple fused images into a video according to the above frame rate, and add the above audio information to obtain a target action video.
[0039] In a second aspect, an embodiment of the present application provides an image processing apparatus, and the apparatus includes:
[0040] A transceiver unit, configured to obtain a target image and determine an action reference video, where the action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object;
[0041] A processing unit, configured to, when it is detected that the target image object is included in the target image and it is determined that the pose of the target image object meets a preset condition, obtain a first 3D model diagram corresponding to the target image object;
[0042] The above processing unit is further configured to obtain a second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video, where the target reference image is one frame of reference image in the multiple frames of reference images;
[0043] The above processing unit is further configured to generate a synthesized image in which the target image object simulates the action of the reference image object in the target reference image according to the target image object, the first 3D model diagram, and the second 3D model diagram;
[0044] Perform video synthesis processing on multiple synthesized images to obtain a target action video.
[0045] Combined with the second aspect, in a possible implementation, the above processing unit is specifically configured to:
[0046] Perform pose recognition processing on the above-mentioned target image to obtain key point information of multiple human key points included in the above-mentioned target image, where the key point information includes position coordinate information and confidence information of the human key points;
[0047] Determine whether the target image includes a target image object based on the above-mentioned multiple human key point information, and determine whether the pose of the target image object meets a preset condition based on the above-mentioned multiple human key point information.
[0048] Combined with the second aspect, in a possible implementation manner, the above-mentioned processing unit is specifically configured to:
[0049] Determine the number of human key points in the above-mentioned multiple human key point information whose confidence information is greater than or equal to a first preset threshold;
[0050] If the number of the above-mentioned human key points is greater than or equal to a second preset threshold, determine that the target image includes a target image object.
[0051] Combined with the second aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include a left shoulder, a left hand, a right shoulder, and a right hand;
[0052] The above-mentioned processing unit is specifically configured to:
[0053] If the abscissa value of the above-mentioned left hand is less than the abscissa value of the above-mentioned left shoulder, the abscissa value of the above-mentioned right hand is greater than the abscissa value of the above-mentioned right shoulder, and the absolute value of the difference between the abscissa value of the above-mentioned left hand and the abscissa value of the above-mentioned right hand is greater than a third preset threshold, determine that the pose of the above-mentioned target image object meets the preset condition;
[0054] Wherein, the above-mentioned third preset threshold is determined by the absolute value of the difference between the abscissa value of the above-mentioned left shoulder and the abscissa value of the above-mentioned right shoulder.
[0055] Combined with the second aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include a left foot and a right foot;
[0056] The above-mentioned processing unit is specifically configured to:
[0057] If the abscissa value of the above-mentioned left foot is less than the abscissa value of the above-mentioned right foot, determine that the pose of the above-mentioned target image object meets the preset condition.
[0058] Combined with the second aspect, in a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include a left eye and a right eye;
[0059] The above-mentioned processing unit is specifically configured to:
[0060] Determine the angle between the line connecting the left eye and the right eye and the horizontal direction according to the position coordinate information of the left eye and the position coordinate information of the right eye described above;
[0061] If the angle is less than the fourth preset threshold, determine that the pose of the target image object meets the preset conditions.
[0062] Combined with the second aspect, in a possible implementation manner, the processing unit is specifically configured to:
[0063] Determine a transformation matrix according to the first 3D model diagram and the second 3D model diagram described above;
[0064] Perform feature extraction processing on the target image object and the first 3D model diagram through a person restoration adversarial neural network, and obtain first image feature parameters output by each layer of the first sampling convolutional network included in the person restoration adversarial neural network, where n is an integer greater than 1;
[0065] Process the transformation matrix, the second 3D model diagram, and the first image feature parameters output by each layer of the first sampling convolutional network through a picture synthesis adversarial neural network to obtain the synthesized image.
[0066] Combined with the second aspect, in a possible implementation manner, the picture synthesis adversarial neural network includes n layers of second sampling convolutional networks, and the n layers of second sampling convolutional networks correspond one-to-one with the n layers of first sampling convolutional networks;
[0067] The processing unit is specifically configured to:
[0068] Input the transformation matrix and the second 3D model diagram into the first layer of the second sampling convolutional network in the n layers of the second sampling convolutional network to obtain second image feature parameters output by the first layer of the second sampling convolutional network;
[0069] Input the second image feature parameters output by the i-th layer of the second sampling convolutional network and the first image feature parameters output by the (i + 1)-th layer of the first sampling convolutional network into the (i + 1)-th layer of the second sampling convolutional network for processing to obtain the second image feature parameters output by the (i + 1)-th layer of the second sampling convolutional network, until the second image feature parameters output by the n-th layer of the second sampling convolutional network are obtained as the synthesized image, where 1 ≤ i ≤ n - 1 and i is an integer.
[0070] Combined with the second aspect, in a possible implementation manner, the processing unit is further configured to:
[0071] Obtain the frame rate and audio information of the action reference video;
[0072] Obtain a background image, and fuse the above-mentioned multiple synthesized images with the above-mentioned background image to generate multiple fused images;
[0073] Synthesize the above-mentioned multiple fused images into a video according to the above-mentioned frame rate, and add the above-mentioned audio information to obtain a target action video.
[0074] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor, a memory, and a transceiver, and the processor, the memory, and the transceiver are connected to each other. The memory is used to store a computer program that supports the terminal device to execute the method provided in the above-mentioned first aspect and / or any possible implementation manner of the first aspect. The computer program includes program instructions, and the processor and the transceiver are configured to call the above-mentioned program instructions to execute the method provided in the above-mentioned first aspect and / or any possible implementation manner of the first aspect.
[0075] In a fourth aspect, an embodiment of the present application provides a server, which includes a processor, a memory, and a transceiver, and the processor, the memory, and the transceiver are connected to each other. The memory is used to store a computer program that supports the server to execute the method provided in the above-mentioned first aspect and / or any possible implementation manner of the first aspect. The computer program includes program instructions, and the processor and the transceiver are configured to call the above-mentioned program instructions to execute the method provided in the above-mentioned first aspect and / or any possible implementation manner of the first aspect.
[0076] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method provided in the above-mentioned first aspect and / or any possible implementation manner of the first aspect.
[0077] In the embodiment of the present application, the server obtains a target image and determines an action reference video. The action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object. When it is detected that the target image object is included in the target image and it is determined that the pose of the target image object meets a preset condition, a first 3D model diagram corresponding to the target image object is obtained. A second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video is obtained. The multiple frames of reference images include the target reference image. According to the target image object, the first 3D model diagram, and the second 3D model diagram, a synthesized image in which the target image object simulates the action of the reference image object in the target reference image is generated. Perform video synthesis processing on the multiple synthesized images to obtain a target action video. By adopting the embodiment of the present application, the user does not need to upload their own dancing video for training the corresponding generator anymore, but only needs to upload a single picture to achieve the generation of a dance video, which improves the image processing efficiency, enhances the user experience, and has high applicability. Brief Description of the Drawings
[0078] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0079] Figure 1 is a schematic diagram of a scenario of image processing provided by an embodiment of the present application;
[0080] Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present application;
[0081] Figure 3 is a schematic diagram of the foreground and background images of a target image provided by an embodiment of the present application;
[0082] Figure 4 is a schematic flowchart of cropping a human image in a target image provided by an embodiment of the present application;
[0083] Figure 5 is a schematic diagram of a target image including various unqualified postures provided by an embodiment of the present application;
[0084] Figure 6 is a schematic diagram of determining a composite image provided by an embodiment of the present application;
[0085] Figure 7 is a schematic diagram of the effect of complementing the background image in a target image provided by an embodiment of the present application;
[0086] Figure 8 is a schematic diagram of a scenario of driving a person in a target image to dance provided by an embodiment of the present application;
[0087] Figure 9 is a schematic diagram of the structure of an image processing device provided by an embodiment of the present application;
[0088] Figure 10 is a schematic diagram of the structure of a network device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0089] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0090] The embodiments of this application relate to Artificial Intelligence (AI) and Machine Learning (ML). Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It mainly produces a new intelligent machine that can respond in a way similar to human intelligence by understanding the essence of intelligence, enabling the intelligent machine to have various functions such as perception, reasoning, and decision-making.
[0091] AI technology is an interdisciplinary subject that mainly includes several major directions such as Computer Vision (CV), speech processing technology, natural language processing technology, and Machine Learning (ML) / Deep Learning. Among them, computer vision technology is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphics processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data; it usually includes technologies such as image processing, video processing, video semantic understanding, and video content / behavior recognition.
[0092] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning or deep learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0093] Based on computer vision technology and machine learning technology in AI technology, an embodiment of the present application provides an image processing method, which includes: obtaining a target image and determining an action reference video, where the action reference video includes multiple reference images, and each reference image in the multiple reference images includes a reference image object. When it is detected that the target image object is included in the target image and it is determined that the pose of the target image object meets a preset condition, obtain a first 3D model image corresponding to the target image object. Obtain a second 3D model image corresponding to the reference image object in the target reference image of the action reference video, and the multiple reference images include the target reference image. Generate a composite image in which the target image object simulates the action of the reference image object in the target reference image, so as to obtain a target action video, that is, perform video synthesis processing on multiple said composite images to obtain a target action video.
[0094] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a scenario of image processing provided by an embodiment of the present application. As Figure 1 shown, the image processing scenario includes a terminal device 101 and a server 102. Among them, the terminal device 101 is the device used by the user, and the terminal device 101 may include, but is not limited to: smart phones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (MID), etc.; the terminal device is configured with a display device, and the display device may also be a monitor, a display screen, a touch screen, etc., and the touch screen may also be a touch panel, a touch panel, etc., which are not limited in the embodiment of the present application.
[0095] The server 102 refers to a background device that can process the target image and the selected action reference video provided by the terminal device 101. After obtaining the target action video based on the target image and the action reference video, the server 102 can return the target action video to the terminal device 101. The server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, multiple servers can also be organized into a blockchain network, and each server is a node in the blockchain network. The terminal device 101 and the server 102 can be directly or indirectly connected through wired communication or wireless communication methods, which are not limited in the present application.
[0096] It should be noted that Figure 1In the image processing scenario shown, the number of terminal devices and servers is only for illustration. For example, the number of terminal devices and servers can be multiple, and the present application does not limit the number of terminal devices and servers. Among them, the method provided by the embodiments of the present application can be applied to servers, and can also be applied to terminal devices, etc., without limitation here. For the convenience of description, the embodiments of the present application may collectively refer to terminal devices and servers as network devices, and the following will be described by taking network devices as an example.
[0097] Next, the method and related devices provided by the embodiments of the present application will be described in detail in conjunction with Figures 2 to 10 respectively.
[0098] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image processing method provided by an embodiment of the present application. The method provided by the embodiments of the present application may include the following steps S201 to S204:
[0099] S201. Obtain a target image and determine an action reference video.
[0100] In some feasible embodiments, the network device obtains a target image and determines an action reference video. Among them, the action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object. The action reference video is a video in which the reference image object performs a certain type of action. Generally speaking, the action reference video can be a dance video, or it can also be a martial arts video, or it can also be a gymnastics video, or it can also be a yoga video, etc., without limitation here. Among them, the embodiments of the present application may take the action reference video as a dance video as an example for description. That is to say, when a user wants to imitate the dance movements of the person in the dance video to be imitated to generate a video of the user dancing this dance, the user can upload a person picture including the user himself, and this person picture is the target image. Then, the server can process the target image and the action reference video uploaded by the user. It can be understood that if the user wants to generate a video of another user (such as the user's friend, the star the user likes) dancing this dance, the target image can also be a person picture including another user, without limitation here.
[0101] Generally speaking, the action reference video can be a video pre-stored in the terminal device or the server. Therefore, the user can select the video to be imitated from the alternative action reference videos such as the action reference video list or from the action reference video library as the action reference video. Optionally, if there is no action reference video that the user wants to imitate among the alternative action reference videos, the user can also upload a video that the user wants to imitate by himself as the action reference video.
[0102] S202. When it is detected that the target image includes a target image object and it is determined that the pose of the target image object meets the preset conditions, obtain the first 3D model diagram corresponding to the target image object.
[0103] In some feasible implementation manners, when it is detected that the target image includes a target image object and it is determined that the pose of the target image object meets the preset conditions, obtain the first 3D model diagram corresponding to the target image object. Generally speaking, the target image can be composed of a target image object (i.e., the foreground) and a background image (which can also be called the background). For example, please refer to Figure 3 , Figure 3 is a schematic diagram of the foreground and background images of the target image provided by the embodiments of the present application. As Figure 3 shown, the image matting technology can be used to process the target image that meets the preset conditions. First, predict the person mask diagram of the target image. According to the person mask diagram, the person and the background image of the target image can be separated. Then, the blank part of the background image after separating the person can be filled in by using the image completion function of OpenMMlab as the background image of the finally generated video.
[0104] Among them, the target image object can be a person image included in the target image, etc., which is not limited here. The first 3D model diagram corresponding to the target image object can be a 3D mesh diagram of the target image object, etc., which is not limited here. That is to say, after the server obtains the target image uploaded by the user, it is also necessary to first determine whether the target image uploaded by the user meets the requirements. Only the target images that meet the requirements can be used for subsequent processing to generate the target action video. For the convenience of description, the following embodiments of the present application will take the target image object as a person image as an example for illustration.
[0105] Specifically, the server can perform pose recognition processing on the target image to obtain key point information of multiple human key points included in the target image. Among them, the key point information includes the position coordinate information and confidence information of the human key points. Therefore, it is possible to determine whether the target image includes a target image object based on the multiple human key point information, and determine whether the pose of the target image object meets the preset conditions based on the multiple human key point information. Among them, the above-mentioned determination of whether the target image includes a target image object based on the multiple human key point information can be understood as: determining the number of human key points in the multiple human key point information whose confidence information is greater than or equal to the first preset threshold. If the number of human key points is greater than or equal to the second preset threshold, it is determined that the target image includes a target image object. For example, the server can process the target image through the OpenPose human pose recognition model to detect whether the user image uploaded by the user includes a person (or a person image, a target person, etc.). Generally speaking, if there is no person in the target image, it means that the target image does not meet the requirements and the process terminates. Optionally, for each frame image in the action reference video (also called a reference image), the same detection can be performed on each frame of the reference image in the action reference video. If there is no person in a certain frame of the reference image, this frame is discarded and the next frame is detected to filter out multiple frame reference images in the action reference video. Each frame of the multiple frame reference images includes a reference image object (i.e., a reference person, or described as a reference figure, a person to be imitated, etc.).
[0106] Among them, if the target image includes a target person, it is further determined whether the pose of the target person included in the target image meets the requirements. Among them, the position coordinate information of each human key point can include the abscissa value and the ordinate value. The multiple human key points can include key points corresponding to the left shoulder, left hand, right shoulder, and right hand respectively. Among them, the above-mentioned determination of whether the pose of the target image object meets the preset conditions based on the multiple human key point information can be understood as: if the abscissa value of the left hand is less than the abscissa value of the left shoulder, the abscissa value of the right hand is greater than the abscissa value of the right shoulder, and the absolute value of the difference between the abscissa value of the left hand and the abscissa value of the right hand is greater than the third preset threshold, it is determined that the pose of the target image object meets the preset conditions. Among them, the third preset threshold is determined by the absolute value of the difference between the abscissa value of the left shoulder and the abscissa value of the right shoulder. For example, the third preset threshold can be equal to the product of the absolute value of the difference between the abscissa value of the left shoulder and the abscissa value of the right shoulder and a preset value. For example, the preset value can be 1.2, etc., and there is no limitation here.
[0107] Optionally, the multiple human key points may include key points corresponding to the left foot and the right foot respectively. Therefore, determining whether the pose of the target image object meets the preset condition according to the multiple human key point information can also be understood as: if the abscissa value of the left foot is less than the abscissa value of the right foot, it is determined that the pose of the target image object meets the preset condition.
[0108] Optionally, the multiple human key points may include key points corresponding to the left eye and the right eye respectively. Therefore, determining whether the pose of the target image object meets the preset condition according to the multiple human key point information can also be understood as: determining the angle between the line connecting the left eye and the right eye and the horizontal direction according to the position coordinate information of the left eye and the position coordinate information of the right eye. If the angle is less than the fourth preset threshold, it is determined that the pose of the target image object meets the preset condition. Wherein, the angle can satisfy: where θ is the angle, d2 represents the absolute value of the difference between the abscissa of the left eye and the abscissa of the right eye, and d1 represents the absolute value of the difference between the ordinate of the left eye and the ordinate of the right eye.
[0109] It can be understood that determining whether the pose of the target image object meets the preset condition according to the multiple human key point information can also be understood as: when the abscissa value of the left hand is less than the abscissa value of the left shoulder, the abscissa value of the right hand is greater than the abscissa value of the right shoulder, and the absolute value of the difference between the abscissa value of the left hand and the abscissa value of the right hand is greater than the third preset threshold, and the abscissa value of the left foot is less than the abscissa value of the right foot, and the angle between the line connecting the left eye and the right eye and the horizontal direction is less than the fourth preset threshold, it is determined that the pose of the target image object meets the preset condition.
[0110] Optionally, in some feasible embodiments, after determining that the target image includes a human image, the human image included in the target image may be cropped first, and then the pose of the cropped human image may be detected. For example, please refer to Figure 4 , Figure 4 is a schematic flowchart of cropping the human image in the target image provided by the embodiment of the present application. As Figure 4As shown, it is assumed that after detecting the target image based on the OpenPose human pose recognition model, the position coordinate information and confidence information corresponding to each of the 25 human key points can be obtained. Then, the maximum and minimum values of the x-axis and y-axis coordinate points in the position coordinate information are taken: min_x, max_x, min_y, max_y. A person box can be formed based on these four points. Centered on this person box, it extends in both the horizontal and vertical axes until one side touches the image edge, and after cropping, a new person box is obtained. If the new person box is a square, an image enlargement or reduction operation is performed to ensure that the final image resolution is 512*512 or 1024*1024; if the new person box is not a square, the image filling operation is performed using the edge pixels of the shorter side to ensure that the person box is a square, and then the image enlargement or reduction operation is performed. Among them, after the above-mentioned cropping operation and image filling operation on the target image, it can be ensured that the person image is in the exact middle of the newly generated target image. Further, for the effect of the final generated video, it is necessary to evaluate the pose of the person image from the following five aspects to determine whether the pose of the person image meets the preset conditions.
[0111] For example, ① the confidence threshold of the confidence information in the human key point information can be set to 0.25. If the confidence of any one of the 25 human key points is less than this threshold, it is considered that the key point is occluded, and this picture is determined to be unqualified, as shown in Figure 5 (a); if all 25 key points are greater than the threshold, it means that all key points in this picture are clearly visible.
[0112] ② Extract the coordinates of the two shoulders and two hands from the 25 human key points, and judge whether the abscissa of the left hand is less than the abscissa of the left shoulder and whether the abscissa of the right hand is greater than the abscissa of the right shoulder. If either one does not meet the condition, it means that at least one hand of the person in the picture is within the shoulder range, and this picture is unqualified, as shown in Figure 5 (b); if both meet the conditions, it means that the two hands of the person in the picture are not within the shoulder range.
[0113] ③ Extract the coordinates of the two shoulders and two hands from the 25 human key points, and judge whether the horizontal distance between the two hands (the absolute value of the difference between the abscissas of the two hand coordinates) is greater than 1.2 times the horizontal distance between the two shoulders. If it does not meet the condition, it means that the two hands of the person in the picture are not opened wide enough, and this picture is unqualified, as shown in Figure 5 (c); if it meets the condition, it means that the two hands of the person in the picture are opened wide enough and do not touch the body.
[0114] ④Extract the coordinates of both feet from 25 human body key points, compare the horizontal coordinates of the left foot with those of the right foot. If the former is greater than the latter, it indicates that there is a crossing phenomenon between the two feet, and this picture is unqualified, as shown in Fig. 5(d); if it is less than or equal to, it means there is no crossing between the two feet.
[0115] ⑤Extract the coordinates of both eyes from 25 human body key points, calculate the horizontal distance between the two eyes and denote it as d1, and the vertical distance between the two eyes as d2. By calculate the angle between the line connecting the two eyes and the horizontal axis. If this angle is greater than 15°, it means that the skew degree of the person's head in the picture is relatively large, and this picture is determined to be unqualified, as Figure 5 (e) shown; if it is less than or equal to 15°, it means that the skew degree of the person in the picture is within the range of acceptable requirements.
[0116] It can be understood that the position coordinates of the key points in each embodiment of the present application are coordinate values within a preset coordinate system. This coordinate system takes the lower left corner of each image as the origin, or this coordinate system can also take the upper left corner of each image as the origin, etc., which is specifically determined according to the application scenario and is not limited here.
[0117] It is understandable that if the target image uploaded by the user meets the evaluation requirements in the above five aspects, it can enter the next step of processing. If any one point is not satisfied, the user needs to make targeted modifications to the uploaded picture. The evaluation of this link is only for the target image, and such evaluation is not performed on each frame image of the action reference video because generally the video has been strictly screened and the evaluation has been done in advance. Optionally, if the action reference video is a video uploaded by the user himself / herself, then for each frame image in the action reference video, the above processing process for the target image also needs to be executed to evaluate the action reference video.
[0118] S203. Obtain the second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video.
[0119] In some feasible implementation manners, to obtain the second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video, multiple frame reference images include the target reference image, that is, the target reference image is one of the multiple frame reference images. That is to say, for the target reference image in the action reference video, the second 3D model diagram corresponding to the reference image object in the target reference image can be obtained, where the second 3D model diagram corresponding to the reference image object can be the 3D mesh diagram of the reference image object, etc., which is not limited here.
[0120] Among them, the embodiments of the present application can process the target image and each frame of the reference image of the action reference video according to the Human Mesh Recovery (HMR) model provided by iPERCore to obtain various parameters characterizing the human postures and shapes in the target image and each frame of the reference image of the action reference video. Furthermore, according to the SMPL model provided by iPERCore, various parameters such as the human postures and shapes of each frame of the image (i.e., the target image and each frame of the reference image in the action reference video) are modeled to respectively model the first initial 3D mesh map corresponding to the target image and the second initial 3D mesh map corresponding to each frame of the reference image in the action reference video. Finally, according to the Neural Mesh Renderer (NMR) model, the initial 3D mesh maps corresponding to each frame of the image output by the SMPL model are deeply rendered to obtain the first 3D mesh map (i.e., the first 3D model map) corresponding to the target image and the second 3D mesh map (i.e., the second 3D model map) corresponding to each frame of the reference image in the action reference video. It can be understood that the HMR model is an end-to-end model for recovering a 3D human model from a 2D image. It is trained based on the two datasets of HumanEva and Huam3.6M and has been open-sourced. The SMPL model is a parametric human model, which is modeled by human shape parameters and pose parameters. That is to say, after inputting the picture into the HMR model for processing, various parameters characterizing the human postures and shapes in the picture can be output. Then, through the SMPL model to model the various parameters output by the HMR model, the initial 3D mesh map can be obtained. The NMR model is a technology for rendering pictures using deep learning, and it is also an open-sourced model. The input of the NMR model can be the initial 3D mesh output by the SMPL model, and the output of the NMR model is the rendered 3D mesh map.
[0121] S204. Generate a composite image in which the target image object simulates the actions of the reference image object in the target reference image, and obtain a target action video based on the composite image.
[0122] In some feasible embodiments, according to the target image object, the first 3D model diagram, and the second 3D model diagram, a composite image in which the target image object simulates the actions of the reference image object in the target reference image can be generated, and then a target action video can be obtained by using the composite image. Specifically, the generation of the composite image in which the target image object simulates the actions of the reference image object in the target reference image according to the target image object, the first 3D model diagram, and the second 3D model diagram can be understood as: determining a transformation matrix according to the first 3D model diagram and the second 3D model diagram. Feature extraction processing is performed on the target image object and the first 3D model diagram through a person restoration adversarial neural network, and first image feature parameters output by each layer of the first sampling convolutional network included in the person restoration adversarial neural network are obtained, where n is an integer greater than 1. The transformation matrix, the second 3D model diagram, and the first image feature parameters output by each layer of the first sampling convolutional network are processed through a picture synthesis adversarial neural network to obtain the composite image.
[0123] Among them, the picture synthesis adversarial neural network includes n layers of second sampling convolutional networks, and the n layers of second sampling convolutional networks correspond one-to-one with the n layers of first sampling convolutional networks. Therefore, the above-mentioned processing of the transformation matrix, the second 3D model diagram, and the first image feature parameters output by each layer of the first sampling convolutional network through the picture synthesis adversarial neural network to obtain the composite image can be understood as: inputting the transformation matrix, the second 3D model diagram, and the first image feature parameters output by the first layer of the first sampling convolutional network in the n layers of first sampling convolutional networks into the first layer of the second sampling convolutional network in the n layers of second sampling convolutional networks to obtain the second image feature parameters output by the first layer of the second sampling convolutional network, inputting the second image feature parameters output by the first layer of the second sampling convolutional network and the first image feature parameters output by the second layer of the first sampling convolutional network into the second layer of the second sampling convolutional network for processing to obtain the second image feature parameters output by the second layer of the second sampling convolutional network; inputting the second image feature parameters output by the second layer of the second sampling convolutional network and the first image feature parameters output by the third layer of the first sampling convolutional network into the third layer of the second sampling convolutional network for processing to obtain the second image feature parameters output by the third layer of the second sampling convolutional network, and so on, inputting the second image feature parameters output by the i-th layer of the second sampling convolutional network and the first image feature parameters output by the (i + 1)-th layer of the first sampling convolutional network into the (i + 1)-th layer of the second sampling convolutional network for processing to obtain the second image feature parameters output by the (i + 1)-th layer of the second sampling convolutional network, until the second image feature parameters output by the n-th layer of the second sampling convolutional network are obtained as the composite image, where 1 ≤ i ≤ n - 1 and i is an integer.
[0124] For easy understanding, please refer toFigure 6 , Figure 6 is a schematic diagram of determining a synthetic image provided by an embodiment of the present application. As Figure 6 shown, taking n = 5 as an example for illustrative purposes, that is, both the human restoration adversarial neural network and the image synthesis adversarial neural network include 5-layer sampling convolutional networks. Among them, the 5-layer sampling convolutional networks included in the human restoration adversarial neural network are respectively the first-layer first sampling convolutional network, the second-layer first sampling convolutional network, the third-layer first sampling convolutional network, the fourth-layer first sampling convolutional network, and the fifth-layer first sampling convolutional network; the 5-layer second sampling convolutional networks included in the image synthesis adversarial neural network are respectively the first-layer second sampling convolutional network, the second-layer second sampling convolutional network, the third-layer second sampling convolutional network, the fourth-layer second sampling convolutional network, and the fifth-layer second sampling convolutional network.
[0125] Among them, as Figure 5 shown, for the human restoration adversarial neural network, input the target image object and the first 3D model diagram into the first-layer first sampling convolutional network of the human restoration adversarial neural network, and the first image feature parameter output by the first-layer first sampling convolutional network can be obtained (that is, F11 as Figure 5 shown). After F11 is processed by the second-layer first sampling convolutional network, the first image feature parameter output by the second-layer first sampling convolutional network can be obtained (that is, F12 as Figure 5 shown). And so on, after layer-by-layer processing, the first image feature parameters output by each subsequent first sampling convolutional network can be obtained, that is, F13, F14, and F15 as Figure 5 shown.
[0126] For the image synthesis adversarial neural network, input the transformation matrix, the second 3D model diagram, and F11 into the first-layer second sampling convolutional network of the image synthesis adversarial neural network, and the second image feature parameter output by the first-layer second sampling convolutional network can be obtained (that is, F21’ as Figure 5 shown). After F21’ and F12 are processed by the second-layer second sampling convolutional network, the second image feature parameter output by the second-layer second sampling convolutional network can be obtained (that is, F22’ as Figure 5 shown). And so on, taking the second image feature parameter output by the i-th layer second sampling convolutional network and the first image feature parameter output by the (i + 1)-th layer first sampling convolutional network as the input of the (i + 1)-th layer second sampling convolutional network for processing, the second image feature parameters output by each subsequent second sampling convolutional network can be obtained, that is, F23’, F24’, and F25’ as Figure 5 shown. Finally, F25’ can be determined as the synthetic image data.
[0127] It should be noted that after obtaining the first 3D mesh map corresponding to the target image object in the target image and the second 3D mesh maps corresponding to the reference image objects included in each frame of the action reference video, the transformation matrix between the target image object and the reference image objects in each frame of the reference image can be further calculated based on the first 3D mesh map and the second 3D mesh maps corresponding to the reference image objects in each frame of the reference image. It can be understood that since the processing of each frame of the reference image in the action reference video is the same, for the convenience of understanding, the following embodiments of the present application will take the processing of one frame of the reference image in the action reference video as an example for illustration, where this one frame of the reference image can be described as the target reference image.
[0128] Among them, when the transformation matrix between the target image object and the reference image object in the target reference image is calculated, the target image object separated by the image matting technology and the first 3D mesh map corresponding to the target image object can be input into the human restoration generative adversarial neural network provided by iPERCore. This human restoration generative adversarial neural network can restore various details of the person (i.e., the target image object) in the target image (such as the texture and color of the clothes, etc.), and these details can be retained in the intermediate layer features of the network. Therefore, the transformation matrix between the target image object and the reference image object in the target reference image and the second 3D mesh map corresponding to the target reference image can be further input into the image synthesis generative adversarial neural network provided by iPERCore, and the above-retained intermediate layer features can be added layer by layer in this network to obtain a synthesized image. It can be understood that the synthesized image is an image of the person in the target image imitating the dance movements of the person in the reference image. Finally, the generated person imitating the dance movements (i.e., the synthesized image) can be added to the background image to generate a fused image. Since the above processing is performed on each frame of the reference image in the action reference video, a series of images of the person in the target image imitating the dance movements of the person in the action reference video can be obtained in this way. Further, these fused pictures can be synthesized into a video according to the frame rate of the action reference video using ffmpeg, and the audio file of the action reference video can be added to generate the final target action video presented to the user (i.e., the video of the person in the target image dancing this dance).
[0129] That is to say, after processing each frame of reference image in the action reference video according to the above steps to generate multiple frames of composite images, the frame rate and audio information of the action reference video can be further obtained, and a background image can be obtained. Then, the multiple composite images are fused with the background image to generate multiple fused images. Finally, the multiple fused images are combined into a video according to the frame rate, and the audio information is added to obtain the target action video. It can be understood that the background image used to generate the fused image in the embodiments of the present application can be the background image included in the target image (such as the filled background shown in Figure 3 ), so that a fused image of the person in the target image performing a dance action on the original background image of the target image can be realized. Optionally, please refer to Figure 7 , Figure 7 which is a schematic diagram of the effect of complementing the background image in the target image provided by the embodiments of the present application. As shown in Figure 7 , if the background image in the target image is too complex, the effect of image complementation will not be very good. Therefore, the background image used to generate the fused image can also be other backgrounds selected by the user, such as a stage, a grassland, etc. Furthermore, a fused image of the person in the target image performing a dance action on a new venue can be realized, increasing the interest.
[0130] Exemplarily, please refer to Figure 8 , Figure 8 which is a schematic diagram of the scene for driving the person in the target image to dance provided by the embodiments of the present application. As shown in Figure 8As shown, when it is determined based on human key point detection that the person in the target image (i.e., the target image object) meets the preset conditions, the target image is cropped to obtain a new person box image. By processing the new person box image through the HMR model + NMR model, a first mesh map corresponding to the target image object can be obtained. And by performing matte processing on the new person box image and complementing the background image obtained by matte processing, the complemented background can be obtained. The second mesh map corresponding to the target reference image in the action reference video is obtained. A transformation matrix can be determined based on the first mesh map and the second mesh map. By performing feature extraction processing on the target image object and the first 3D model map through the person restoration adversarial neural network, the first image feature parameters (i.e., intermediate layer features) output by each of the n layers of the first sampling convolutional network included in the person restoration adversarial neural network can be obtained. By processing the transformation matrix, the second 3D model map, and the intermediate layer features through the image synthesis adversarial neural network, a synthesized image in which the person in the target image simulates the action of the person in the target reference image can be obtained. Finally, by performing fusion processing on the synthesized image and the above-mentioned complemented background, or by performing fusion processing on the synthesized image and a new background, a fused image can be obtained. Therefore, a target action video can be generated based on multiple fused images to realize the generation of the dance action of the person in the target image.
[0131] In the embodiment of the present application, the server obtains the target image and determines the action reference video. The action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object. When the server detects that the target image includes the target image object and determines that the pose of the target image object meets the preset conditions, the first 3D model map corresponding to the target image object is obtained. The second 3D model map corresponding to the reference image object in the target reference image of the action reference video is obtained, and the multiple frames of reference images include the target reference image. According to the target image object, the first 3D model map, and the second 3D model map, a synthesized image in which the target image object simulates the action of the reference image object in the target reference image is generated, so as to obtain the target action video, that is, video synthesis processing is performed on multiple synthesized images to obtain the target action video. By adopting the embodiment of the present application, the image processing efficiency can be improved, the user experience can be enhanced, and the applicability is high.
[0132] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the image processing device provided by the embodiment of the present application. The image processing device provided by the embodiment of the present application includes:
[0133] A transceiver unit 91, configured to obtain the target image and determine the action reference video, where the action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object;
[0134] A processing unit 92, configured to obtain a first 3D model diagram corresponding to the target image object when it is detected that the target image object is included in the above target image and it is determined that the pose of the target image object meets a preset condition;
[0135] The above processing unit 92 is further configured to obtain a second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video, where the target reference image is one of the multiple reference images;
[0136] The above processing unit 92 is further configured to generate a composite image in which the target image object simulates the action of the reference image object in the target reference image according to the target image object, the first 3D model diagram, and the second 3D model diagram;
[0137] Perform video composition processing on multiple said composite images to obtain a target action video.
[0138] In a possible implementation manner, the above processing unit 92 is specifically configured to:
[0139] Perform pose recognition processing on the above target image to obtain key point information of multiple human key points included in the above target image, where the key point information includes position coordinate information and confidence information of the human key points;
[0140] Determine whether the target image object is included in the above target image according to the above multiple human key point information, and determine whether the pose of the target image object meets a preset condition according to the above multiple human key point information.
[0141] In a possible implementation manner, the above processing unit 92 is specifically configured to:
[0142] Determine the number of human key points in the above multiple human key point information whose confidence information is greater than or equal to a first preset threshold;
[0143] If the number of the above human key points is greater than or equal to a second preset threshold, it is determined that the target image object is included in the above target image.
[0144] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above multiple human key points include a left shoulder, a left hand, a right shoulder, and a right hand;
[0145] The above processing unit 92 is specifically configured to:
[0146] If the abscissa value of the left hand is less than the abscissa value of the left shoulder, the abscissa value of the right hand is greater than the abscissa value of the right shoulder, and the absolute value of the difference between the abscissa value of the left hand and the abscissa value of the right hand is greater than a third preset threshold, it is determined that the pose of the target image object meets the preset condition;
[0147] Wherein, the third preset threshold is determined by the absolute value of the difference between the abscissa value of the left shoulder and the abscissa value of the right shoulder.
[0148] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the multiple human key points include a left foot and a right foot;
[0149] The processing unit 92 is specifically configured to:
[0150] If the abscissa value of the left foot is less than the abscissa value of the right foot, it is determined that the pose of the target image object meets the preset condition.
[0151] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the multiple human key points include a left eye and a right eye;
[0152] The processing unit 92 is specifically configured to:
[0153] Determine the angle between the line connecting the left eye and the right eye and the horizontal direction according to the position coordinate information of the left eye and the position coordinate information of the right eye;
[0154] If the angle is less than a fourth preset threshold, it is determined that the pose of the target image object meets the preset condition.
[0155] In a possible implementation manner, the processing unit 92 is specifically configured to:
[0156] Determine a transformation matrix according to the first 3D model diagram and the second 3D model diagram;
[0157] Perform feature extraction processing on the target image object and the first 3D model diagram through a person restoration adversarial neural network, and obtain first image feature parameters output by each of the n first sampling convolutional networks included in the person restoration adversarial neural network, where n is an integer greater than 1;
[0158] Process the transformation matrix, the second 3D model diagram, and the first image feature parameters output by each layer of the first sampling convolutional network through a picture synthesis adversarial neural network to obtain the synthesized image.
[0159] In a possible implementation manner, the above-mentioned picture synthesis adversarial neural network includes n layers of second sampling convolutional networks, and the n layers of second sampling convolutional networks correspond one-to-one with the n layers of first sampling convolutional networks;
[0160] The above-mentioned processing unit 92 is specifically configured to:
[0161] Input the above-mentioned transformation matrix and the second 3D model diagram into the first layer of the second sampling convolutional network among the n layers of second sampling convolutional networks, and obtain the second image feature parameters output by the first layer of the second sampling convolutional network;
[0162] Input the second image feature parameters output by the i-th layer of the second sampling convolutional network and the first image feature parameters output by the (i + 1)-th layer of the first sampling convolutional network into the (i + 1)-th layer of the second sampling convolutional network for processing, and obtain the second image feature parameters output by the (i + 1)-th layer of the second sampling convolutional network until the second image feature parameters output by the n-th layer of the second sampling convolutional network are obtained as the synthesized image, where 1 ≤ i ≤ n - 1 and i is an integer.
[0163] In a possible implementation manner, the above-mentioned processing unit 92 is further configured to:
[0164] Obtain the frame rate and audio information of the above-mentioned action reference video;
[0165] Obtain a background picture, and fuse the above-mentioned multiple synthesized images with the background picture to generate multiple fused images;
[0166] Synthesize the above-mentioned multiple fused images into a video according to the above-mentioned frame rate, and add the above-mentioned audio information to obtain a target action video.
[0167] In the embodiments of the present application, the image processing device can obtain a target image and determine an action reference video. The action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object. When it is detected that the target image object is included in the target image and it is determined that the posture of the target image object meets a preset condition, the first 3D model diagram corresponding to the target image object is obtained. The second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video is obtained, and the multiple frames of reference images include the target reference image. According to the target image object, the first 3D model diagram, and the second 3D model diagram, a synthesized image in which the target image object simulates the action of the reference image object in the target reference image is generated, so as to obtain a target action video, that is, the multiple synthesized images are subjected to video synthesis processing to obtain a target action video. By adopting the embodiments of the present application, the image processing efficiency can be improved, the user experience can be enhanced, and the applicability is high.
[0168] Please refer to Figure 10 , Figure 10It is a schematic structural diagram of a network device provided by an embodiment of the present application. As shown in FIG. 10, the network device in this embodiment may include: one or more processors 1001, a memory 1002, and a transceiver 1003. The above-mentioned processors 1001, memory 1002, and transceiver 1003 are connected through a bus 1004. The memory 1002 is used to store a computer program, and the computer program includes program instructions. The processors 1001 and transceiver 1003 are used to execute the program instructions stored in the memory 1002 and perform the following operations:
[0169] The transceiver 1003 is used to obtain a target image and determine an action reference video, where the action reference video includes multiple reference images, and each reference image in the multiple reference images includes a reference image object;
[0170] The processor 1001 is used to obtain a first 3D model diagram corresponding to the target image object when it is detected that the target image object is included in the target image and it is determined that the pose of the target image object meets a preset condition;
[0171] The processor 1001 is used to obtain a second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video, where the target reference image is one of the multiple reference images;
[0172] The processor 1001 is used to generate a composite image in which the target image object simulates the action of the reference image object in the target reference image according to the target image object, the first 3D model diagram, and the second 3D model diagram;
[0173] The processor 1001 is used to perform video composition processing on multiple composite images to obtain a target action video.
[0174] In a possible implementation manner, the processor 1001 is further used to:
[0175] Perform pose recognition processing on the target image to obtain key point information of multiple human key points included in the target image, where the key point information includes position coordinate information and confidence information of the human key points;
[0176] Determine whether the target image object is included in the target image according to the multiple human key point information, and determine whether the pose of the target image object meets the preset condition according to the multiple human key point information.
[0177] In a possible implementation manner, the processor 1001 is further used to:
[0178] Determine the number of human key points in the above-mentioned multiple human key point information whose confidence information is greater than or equal to the first preset threshold;
[0179] If the number of the above-mentioned human key points is greater than or equal to the second preset threshold, it is determined that the target image includes a target image object.
[0180] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include the left shoulder, the left hand, the right shoulder, and the right hand;
[0181] The above-mentioned processor 1001 is further configured to:
[0182] If the abscissa value of the above-mentioned left hand is less than the abscissa value of the above-mentioned left shoulder, the abscissa value of the above-mentioned right hand is greater than the abscissa value of the above-mentioned right shoulder, and the absolute value of the difference between the abscissa value of the above-mentioned left hand and the abscissa value of the above-mentioned right hand is greater than the third preset threshold, it is determined that the posture of the above-mentioned target image object meets the preset condition;
[0183] Wherein, the above-mentioned third preset threshold is determined by the absolute value of the difference between the abscissa value of the above-mentioned left shoulder and the abscissa value of the above-mentioned right shoulder.
[0184] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include the left foot and the right foot;
[0185] The above-mentioned processor 1001 is further configured to:
[0186] If the abscissa value of the above-mentioned left foot is less than the abscissa value of the above-mentioned right foot, it is determined that the posture of the above-mentioned target image object meets the preset condition.
[0187] In a possible implementation manner, the position coordinate information includes an abscissa value and an ordinate value; the above-mentioned multiple human key points include the left eye and the right eye;
[0188] The above-mentioned processor 1001 is further configured to:
[0189] Determine the included angle between the line connecting the above-mentioned left eye and the above-mentioned right eye and the horizontal direction according to the position coordinate information of the above-mentioned left eye and the position coordinate information of the above-mentioned right eye;
[0190] If the above-mentioned included angle is less than the fourth preset threshold, it is determined that the posture of the above-mentioned target image object meets the preset condition.
[0191] In a possible implementation manner, the above-mentioned processor 1001 is further configured to:
[0192] Determine a transformation matrix according to the above-mentioned first 3D model diagram and the above-mentioned second 3D model diagram;
[0193] The above-mentioned target image object and the above-mentioned first 3D model diagram are subjected to feature extraction processing by a person restoration confrontation neural network, and first image feature parameters output by each layer of the first sampling convolutional network included in the above-mentioned person restoration confrontation neural network are obtained, where n is an integer greater than 1;
[0194] The above-mentioned conversion matrix, the above-mentioned second 3D model diagram, and the first image feature parameters output by each layer of the first sampling convolutional network are processed by an image synthesis confrontation neural network to obtain the above-mentioned synthesized image.
[0195] In a possible implementation manner, the above-mentioned image synthesis confrontation neural network includes n layers of second sampling convolutional networks, and the n layers of second sampling convolutional networks correspond one-to-one with the n layers of first sampling convolutional networks;
[0196] The above-mentioned processor 1001 is further configured to:
[0197] Input the above-mentioned conversion matrix and the above-mentioned second 3D model diagram into the first layer of the second sampling convolutional network among the n layers of second sampling convolutional networks to obtain second image feature parameters output by the first layer of the second sampling convolutional network;
[0198] Input the second image feature parameters output by the i-th layer of the second sampling convolutional network and the first image feature parameters output by the (i + 1)-th layer of the first sampling convolutional network into the (i + 1)-th layer of the second sampling convolutional network for processing to obtain the second image feature parameters output by the (i + 1)-th layer of the second sampling convolutional network, until the second image feature parameters output by the n-th layer of the second sampling convolutional network are obtained as the above-mentioned synthesized image, where 1 ≤ i ≤ n - 1 and i is an integer.
[0199] In a possible implementation manner, the above-mentioned processor 1001 is further configured to:
[0200] Obtain the frame rate and audio information of the above-mentioned action reference video;
[0201] Obtain a background picture, and fuse the above-mentioned multiple synthesized images with the background picture to generate multiple fused images;
[0202] Synthesize the above-mentioned multiple fused images into a video according to the above-mentioned frame rate, and add the above-mentioned audio information to obtain a target action video.
[0203] It should be understood that in some feasible embodiments, the above-mentioned processor 1001 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory 1002 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1001. A part of the memory 1002 may also include a non-volatile random access memory. For example, the memory 1002 may also store information about the device type.
[0204] In specific implementation, the above network device may execute, through each of its built-in functional modules, the implementation manners provided in each of the steps as described above Figure 2 to Figure 8 in the implementation manners provided in each of the steps. For specific details, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.
[0205] In the embodiments of the present application, the network device may acquire a target image and determine an action reference video. The action reference video includes multiple frames of reference images, and each frame of reference image in the multiple frames of reference images includes a reference image object. When it is detected that the target image object is included in the target image and it is determined that the pose of the target image object meets a preset condition, a first 3D model diagram corresponding to the target image object is acquired. A second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video is acquired. The multiple frames of reference images include the target reference image. According to the target image object, the first 3D model diagram, and the second 3D model diagram, a composite image in which the target image object simulates the action of the reference image object in the target reference image is generated, so as to obtain a target action video, that is, video synthesis processing is performed on multiple said composite images to obtain the target action video. By adopting the embodiments of the present application, the image processing efficiency can be improved, the user experience can be enhanced, and the applicability is high.
[0206] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the Figures 2 to 8 image processing method provided in each of the steps is implemented. For specific details, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.
[0207] The above computer-readable storage medium may be the internal storage unit of the image processing device provided in any of the foregoing embodiments or the above terminal device, such as the hard disk or memory of an electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.
[0208] The terms "first", "second", "third", "fourth", etc. in the claims, the description and the drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0209] Referring to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase presented at various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0210] The method and related devices provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining a target image and determining an action reference video, where the action reference video includes multiple reference images, and each reference image in the multiple reference images includes a reference image object; Performing pose recognition processing on the target image to obtain key point information of multiple human key points included in the target image, where the key point information includes position coordinate information and confidence information of the human key points; Determining whether the target image includes a target image object according to the multiple human key point information, and determining whether the pose of the target image object meets a preset condition according to the multiple human key point information; When it is detected that the target image includes the target image object and it is determined that the pose of the target image object meets the preset condition, obtaining a first 3D model diagram corresponding to the target image object; Obtaining a second 3D model diagram corresponding to the reference image object in the target reference image of the action reference video, where the target reference image is one of the multiple reference images; Determining a transformation matrix according to the first 3D model diagram and the second 3D model diagram; Performing feature extraction processing on the target image object and the first 3D model diagram through a human restoration adversarial neural network, and obtaining first image feature parameters output by each of the n first sampling convolutional networks included in the human restoration adversarial neural network, where n is an integer greater than 1; Inputting the transformation matrix, the second 3D model diagram, and the first image feature parameters output by the first layer of the first sampling convolutional network into the first layer of the second sampling convolutional network included in the picture synthesis adversarial neural network to obtain second image feature parameters output by the first layer of the second sampling convolutional network; the n layers of the second sampling convolutional network correspond one-to-one with the n layers of the first sampling convolutional network; Inputting the second image feature parameters output by the i-th layer of the second sampling convolutional network and the first image feature parameters output by the (i + 1)-th layer of the first sampling convolutional network into the (i + 1)-th layer of the second sampling convolutional network for processing to obtain second image feature parameters output by the (i + 1)-th layer of the second sampling convolutional network until the second image feature parameters output by the n-th layer of the second sampling convolutional network are obtained, where 1 ≤ i ≤ n - 1 and i is an integer; Taking the second image feature parameters output by the n-th layer of the second sampling convolutional network as a synthesized image of the target image object simulating the action of the reference image object in the target reference image; Performing video synthesis processing on multiple of the synthesized images to obtain a target action video.
2. The method according to claim 1, wherein The determining whether the target image includes a target image object according to the multiple human key point information includes: Determining the number of human key points in the multiple human key point information whose confidence information is greater than or equal to a first preset threshold; If the number of human key points is greater than or equal to a second preset threshold, determining that the target image includes a target image object.
3. The method according to claim 1 or 2, characterized in that, The position coordinate information includes an abscissa value and an ordinate value; the multiple human key points include a left shoulder, a left hand, a right shoulder, and a right hand. Determining whether the pose of the target image object meets a preset condition according to the multiple human key point information includes: If the abscissa value of the left hand is less than the abscissa value of the left shoulder, the abscissa value of the right hand is greater than the abscissa value of the right shoulder, and the absolute value of the difference between the abscissa value of the left hand and the abscissa value of the right hand is greater than a third preset threshold, it is determined that the pose of the target image object meets the preset condition; Wherein, the third preset threshold is determined by the absolute value of the difference between the abscissa value of the left shoulder and the abscissa value of the right shoulder.
4. The method according to claim 1 or 2, characterized in that, The position coordinate information includes an abscissa value and an ordinate value; The multiple human key points include the left foot and the right foot; Determining whether the pose of the target image object meets a preset condition according to the multiple human key point information includes: If the abscissa value of the left foot is less than the abscissa value of the right foot, it is determined that the pose of the target image object meets the preset condition.
5. The method according to claim 1 or 2, characterized in that, The position coordinate information includes an abscissa value and an ordinate value; The multiple human key points include the left eye and the right eye; Determining whether the pose of the target image object meets a preset condition according to the multiple human key point information includes: Determining the angle between the line connecting the left eye and the right eye and the horizontal direction according to the position coordinate information of the left eye and the position coordinate information of the right eye; If the angle is less than a fourth preset threshold, it is determined that the pose of the target image object meets the preset condition.
6. The method according to claim 1 or 2, characterized in that, Performing video composition processing on the multiple synthesized images to obtain a target action video includes: Obtaining the frame rate and audio information of the action reference video; Obtaining a background picture, and fusing the multiple synthesized images with the background picture to generate multiple fused images; Synthesizing the multiple fused images into a video according to the frame rate, and adding the audio information to obtain a target action video.
7. A terminal device, characterized in that, Including a processor, a memory, and a transceiver, the processor, the memory, and the transceiver are connected to each other; The memory is used to store a computer program, the computer program includes program instructions, and the processor and the transceiver are configured to call the program instructions to execute the method according to any one of claims 1-6.
8. A server, characterized in that, Including a processor, a memory, and a transceiver, the processor, the memory, and the transceiver are connected to each other; The memory is used to store a computer program, the computer program includes program instructions, and the processor and the transceiver are configured to call the program instructions to execute the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 1-6.
Citation Information
Patent Citations
Real person video generation method and device, readable storage medium and equipment
CN112613495A