Image processing model training method and device, equipment and storage medium
By using the first sample set and the second sample set in the training of the image processing model and adjusting the model parameters using inter-frame loss, the inter-frame jitter problem during object replacement in the video clip is solved, and the image quality is improved.
Patent Information
- Application Number
- CN202311545098.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, the image processing model is prone to inter-frame jitter when object replacement in video clips, resulting in unstable image quality.
By obtaining the first sample set and the second sample set, the image processing model to be trained is iteratively trained, and the model parameters are adjusted using inter-frame loss to ensure inter-frame stability.
It effectively avoids inter-frame jitter, improves the image quality of the processed video images, and is suitable for image processing in complex scenarios.
Smart Images

Figure CN120020902A_ABST
Abstract
Description
Background Art
[0002] In application scenarios such as film and television production, game design, virtual avatars, and privacy protection, it is usually adopted that one object replaces another object in an image to speed up the production of multimedia information, or avoid the appearance of sensitive and illegal information, or avoid the leakage of privacy data, etc.
[0003] In order to improve the processing efficiency, an image processing model is usually adopted to implement the replacement operation of the object in the image. However, under the related technologies, the image processing models are all trained by using a single image and label information as training samples, that is, the reference information in the training samples is not comprehensive enough. In this way, when the object in each video image included in the video segment is replaced by the image processing model, inter-frame jitter will occur, so that the inter-frame stability cannot be maintained, and thus the image quality of the processed video image cannot be guaranteed.
[0004] Therefore, how to ensure the accuracy of the image processing model and improve the image quality of the processed video image is a technical problem that needs to be solved currently. Summary of the Invention
[0005] The embodiments of the present application provide a training method, device, equipment and storage medium for an image processing model to ensure the accuracy of the image processing model and improve the image quality of the processed video image.
[0006] In a first aspect, the embodiments of the present application provide a training method for an image processing model, and the method includes:
[0007] Obtain a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein, the preset replacement image includes: a replacement object for replacing a target object in the first training image;
[0008] For each first sample, based on a pre-constructed target occlusion image, perform occlusion processing on the first training image in the first sample to obtain a corresponding second training image, and construct a second sample associated with each first sample based on the second training image and the preset replacement image in the first sample;
[0009] Based on the first sample set and the second sample set, iteratively train the image processing model to be trained to obtain the target image processing model, where the second sample set includes the second samples associated with each first sample; in one round of iterative training, adjust the parameters of the image processing model to be trained based on the inter-frame loss determined by the first predicted image and the second predicted image; the first predicted image is obtained by processing the corresponding first training image through the image processing model to be trained based on the preset replacement image in the selected first sample; the second predicted image is obtained by processing the corresponding second training image through the image processing model to be trained based on the preset replacement image in the second sample associated with the selected first sample.
[0010] In a second aspect, an embodiment of the present application provides a training device for an image processing model, and the device includes:
[0011] An acquisition unit, configured to acquire a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein, the preset replacement image includes: a replacement object for replacing the target object in the first training image;
[0012] A processing unit, configured to, for each first sample, perform occlusion processing on the first training image in the first sample based on a pre-constructed target occlusion image to obtain a corresponding second training image, and construct a second sample associated with each first sample based on the second training image and the preset replacement image in the first sample;
[0013] A training unit, configured to iteratively train the image processing model to be trained based on the first sample set and the second sample set to obtain the target image processing model, where the second sample set includes the second samples associated with each first sample; in one round of iterative training, adjust the parameters of the image processing model to be trained based on the inter-frame loss determined by the first predicted image and the second predicted image; the first predicted image is obtained by processing the corresponding first training image through the image processing model to be trained based on the preset replacement image in the selected first sample; the second predicted image is obtained by processing the corresponding second training image through the image processing model to be trained based on the preset replacement image in the second sample associated with the selected first sample.
[0014] In a possible implementation manner, the target occlusion image includes an occlusion object; the processing unit is specifically configured to:
[0015] Perform occlusion processing on the first training image based on the mask image of the target occlusion image to obtain a masked training image; the target area in the masked training image that matches the occlusion object is occluded; the mask image is used to determine the position area of the occlusion object in the target occlusion image;
[0016] Based on the mask image of the target occluded image, perform occlusion processing on the target occluded image to obtain a masked occluded image; other regions in the masked occluded image except for the target region matching the occluding object are occluded;
[0017] Overlay the masked training image and the masked occluded image to obtain a second training image.
[0018] In a possible implementation, the training unit is specifically configured to:
[0019] In one round of iterative training, perform the following operations:
[0020] Select a first sample from the first sample set and a second sample associated with the first sample from the second sample set; wherein, the first sample and the second sample contain the same preset replacement image;
[0021] Input the first sample and the second sample into the image processing model to be trained to obtain a first predicted image and a second predicted image;
[0022] Based on the first predicted image and the second predicted image, construct a target loss function; wherein, the target loss function includes an inter-frame loss;
[0023] Based on the target loss function, adjust the parameters of the image processing model to be trained.
[0024] In a possible implementation, the inter-frame loss includes an inter-frame feature loss; and the inter-frame feature loss is determined by the following method:
[0025] Input the first predicted image and the second predicted image into a trained feature extraction network; wherein, the feature extraction network includes at least one network layer;
[0026] In each network layer, determine the first image feature of the first predicted image and the second image feature of the second predicted image;
[0027] Determine the feature difference between the first image feature and the second image feature, and based on at least one feature difference, determine the inter-frame feature loss.
[0028] In a possible implementation, the inter-frame loss includes an inter-frame image loss; and the inter-frame image loss is determined by the following method:
[0029] Based on the first predicted image and the mask image of the target occluded image, determine the first non-occluded region in the first predicted image;
[0030] Based on the second predicted image and the mask image of the target occluded image, determine the second non-occluded region in the second predicted image;
[0031] Determine the inter-frame image loss based on the regional image difference between the first non-occluded region and the second non-occluded region.
[0032] In a possible implementation, the first sample further includes a first labeled image, and the second sample further includes a second labeled image, where the second labeled image is obtained by performing an occlusion process on the first labeled image based on the target occlusion image;
[0033] The target loss function further includes: a prediction loss determined based on the difference between the first predicted image and the first labeled image, or a prediction loss determined based on the difference between the second predicted image and the second labeled image; the prediction loss includes: a predicted image loss and a predicted feature loss.
[0034] In a possible implementation, the target loss function further includes an image generation loss; and the image generation loss is determined in the following manner:
[0035] Determine the image generation loss based on the similarity between the first predicted image and a preset replacement image; or determine the image generation loss based on the similarity between the second predicted image and a preset replacement image.
[0036] In a possible implementation, the target loss function further includes an adversarial loss; and the adversarial loss is determined in the following manner:
[0037] Input the first predicted image and the first labeled image in the first sample into a trained discriminative network to obtain the discriminative result of the first predicted image, and determine the adversarial loss based on the discriminative result; where the discriminative result is used to characterize whether the first predicted image is a real image; or
[0038] Input the second predicted image and the second labeled image in the second sample into a trained discriminative network to obtain the discriminative result of the second predicted image, and determine the adversarial loss based on the discriminative result; where the discriminative result is used to characterize whether the second predicted image is a real image.
[0039] In a possible implementation, the target occlusion image is obtained in the following manner:
[0040] Obtain a to-be-processed occlusion image, where the to-be-processed occlusion image includes at least an occlusion object;
[0041] Perform a segmentation process on the to-be-processed occlusion image through a trained segmentation network to obtain the target occlusion image and an associated mask image; the target occlusion image includes the occlusion object, and the mask image is used to determine the position of the occlusion object in the target occlusion image.
[0042] In a possible implementation, the video image to be trained and the first preset replacement image are obtained after image preprocessing; the video image to be trained contains a target object, and the first preset replacement image contains a replacement object.
[0043] In a possible implementation, after obtaining the target image processing model, the training unit is further configured to:
[0044] Based on the replacement object in the target replacement image, the replacement object in the video image to be processed is replaced through the target image processing model to obtain a target video image.
[0045] In a third aspect, an embodiment of the present application provides a computing device, including: a memory and a processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the image processing model training method provided by the embodiment of the present application.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the image processing model training method provided by the embodiment of the present application are implemented.
[0047] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; when the processor of the computing device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computing device executes the steps of the image processing model training method provided by the embodiment of the present application.
[0048] The beneficial effects of the present application are as follows:
[0049] An embodiment of the present application provides a method, device, equipment and storage medium for training an image processing model, which relates to technical fields such as image processing and artificial intelligence, and can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0050] In the embodiment of the present application, considering the problem that when training an image processing model based on a single sample set in the related art, due to the lack of comprehensive reference information of a single sample, the training accuracy of the model is low, an implementation method of using two associated sample sets for training in one round of iterative training is proposed; that is, based on the first sample set and the second sample set, the image processing model to be trained is iteratively trained to obtain a target image processing model.
[0051] When iteratively training a training image processing model, first obtain a first sample set; each first sample in the first sample set includes at least a video image to be trained and a first preset replacement image, and the first preset replacement image contains a replacement object for replacing the target object in the video image to be trained; then, based on a pre-constructed target occlusion image, perform occlusion on the first preset replacement image to obtain a corresponding second preset replacement image, and construct a second sample based on the second preset replacement image and the video image to be trained associated with the first preset replacement image; therefore, for each first sample, a second sample associated with it can be determined, and a second sample set is obtained based on the second sample. By constructing the first sample and the associated second sample, the scenario of object occlusion during frame movement in a video can be simulated.
[0052] After obtaining the first sample set and the second sample set, based on the first sample set and the second sample set, iteratively train the image processing model to be trained to obtain a target image processing model; in one round of iterative training: first, select a first sample from the first sample set and a second sample associated with the first sample from the second sample set; then, input the first preset replacement image and the video image to be trained in the first sample into the image processing model to be trained, and through the image processing model to be trained, process the video image to be trained based on the first preset replacement image to obtain a first predicted image, and input the second preset replacement image and the video image to be trained in the second sample into the image processing model to be trained, and through the image processing model to be trained, process the video image to be trained based on the second preset replacement image to obtain a second predicted image; finally, adjust the parameters of the image processing model to be trained based on the inter-frame loss determined by the first predicted image and the second predicted image. By using the inter-frame loss to constrain the inter-frame stability of the video segment, when replacing objects in each video image included in the video segment in the case of a moving object occluding the object, the problem of inter-frame jitter can be prevented.
[0053] In summary, the target image processing model obtained by training through the model training method provided in the embodiments of the present application can maintain inter-frame stability when replacing objects in each video image in a video segment, thereby ensuring the image quality of the processed video images. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0056] Figure 2 Schematic diagram of an application example provided by an embodiment of the present application;
[0057] Figure 3 Model structure diagram of an image processing model provided by an embodiment of the present application;
[0058] Figure 4 Flowchart of a method for constructing a first sample provided by an embodiment of the present application;
[0059] Figure 5 Schematic diagram of a first sample example provided by an embodiment of the present application;
[0060] Figure 6 Schematic diagram of constructing a first training image provided by an embodiment of the present application;
[0061] Figure 7 Schematic diagram of constructing a second sample based on a target occlusion image provided by an embodiment of the present application;
[0062] Figure 8 Flowchart of a method for constructing a target occlusion image provided by an embodiment of the present application;
[0063] Figure 9 Schematic diagram of constructing a target occlusion image provided by an embodiment of the present application;
[0064] Figure 10 Flowchart of a training method for an image processing model provided by an embodiment of the present application;
[0065] Figure 11 Schematic diagram of extracting features through a feature extraction network provided by an embodiment of the present application;
[0066] Figure 12 Schematic diagram of training an image processing model provided by an embodiment of the present application;
[0067] Figure 13 Structure diagram of a training device for an image processing model provided by an embodiment of the present application;
[0068] Figure 14 Structure diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0069] In order to make the objectives, technical solutions and beneficial effects of the present application clearer and more understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0070] The following explains some terms in the embodiments of the present application to facilitate the understanding of those skilled in the art.
[0071] Adaptive Instance Normalization (AdaIN): Also known as an image style transfer algorithm, it is to transfer the style, texture, etc. in a style image to another content image while retaining the main structure of the content image. For example, both the style image and the content image are face images, and the style image contains the face image of person A, and the content image contains the face image of person B. At this time, the facial features of person A in the style image are transferred to the content image, and the original expression, angle, background, etc. information in the content image is still retained. The image processing process in the embodiments of the present application can be understood as a process of image style transfer, that is, the replacement object in the target replacement image is transferred to the video image to be processed through an image processing model, and the replacement object in the video image to be processed is replaced with the replacement object, and the generated target video image still retains the expression, angle, scene, etc. information of the person in the video image to be processed.
[0072] Learned Perceptual Image Patch Similarity (LPIPS) loss: Also known as perceptual loss, it is used to measure the difference between two images. The lower the value of LPIPS, the more similar the two images are, and vice versa, the greater the difference.
[0073] Cosine Similarity, also known as cosine similarity, measures the similarity between two vectors by calculating the cosine value of the angle between the two vectors. For example, in the embodiments of the present application, after obtaining two predicted images, the image feature vectors of the two predicted images are respectively determined, and then the similarity is measured based on the cosine values of the angles between the two image feature vectors and the feature vector of the preset replacement image.
[0074] The term "exemplary" used hereinafter means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" does not have to be construed as superior to or better than other embodiments.
[0075] The terms "first" and "second" in the text are only for descriptive purposes and should not be construed as indicating explicitly or implicitly relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0076] In application scenarios such as film and television production, game design, virtual avatars, and privacy protection, it is usually adopted to replace one object in an image with another object; to speed up the production of multimedia information, or to avoid the appearance of sensitive or illegal information, or to avoid the leakage of privacy data, etc. For example: in film and television production, when an actor is unable to complete a professional action, a professional can complete the recording. Later, for the recorded video clip, the facial features of the actor are used to replace the facial features of the professional in each video image in the video clip, so that when the viewer watches this video clip, it seems that the actor himself has completed it. Or, to avoid re-shooting the video to save costs, when it is determined that an actor corresponding to a certain character in a filmed movie or TV drama needs to be replaced, the face-changing operation can also be performed by means of image replacement. Or, in virtual avatars, virtual characters can be used for face-changing to improve the fun and privacy protection of the live broadcast.
[0077] In order to improve the processing efficiency and accuracy, an image processing model is usually adopted to implement the object replacement operation for the image. However, under the related technologies, the image processing models are all trained with a single image and label information as training samples, that is, the reference information in the training samples is not comprehensive enough. In this way, when the object replacement is performed on each video image included in the video clip through the image processing model, frame jitter will occur, so that the inter-frame stability cannot be maintained, and thus the image quality of the processed video image cannot be guaranteed. For example: in a film and television production site, it is a relatively common scene for moving objects such as scattered snowflakes, leaves, and flower petals. In these scenes, it is easy for moving objects to move on the face or body of a film and television character (for example, in a video clip, the face of the film and television character basically remains motionless, but the moving object moves on the face). At this time, when the object replacement is performed on each video image included in the video clip by using the image processing model in the related technologies, frame jitter will occur and the inter-frame stability under the movement of the object cannot be maintained.
[0078] Therefore, how to ensure the accuracy of the image processing model and improve the image quality of the processed video image is a technical problem that needs to be solved currently.
[0079] In view of this, in the embodiments of the present application, considering the inter-frame stability under the movement of an object and in order to prevent the problem of inter-frame jitter, a training method, device, equipment and storage medium for an image processing model are proposed to ensure the accuracy of the image processing model, thereby improving the image quality of the video image processed by the image processing model and meeting the image processing in complex scenarios with moving objects such as film and television production at the same time.
[0080] In the embodiments of the present application, considering the problem that the accuracy of model training is low due to the insufficient comprehensive reference information of a single sample when training an image processing model based on a single sample set in the related art, an implementation manner of using two associated sample sets for training in one round of iterative training is proposed; that is, based on a first sample set and a second sample set, the image processing model to be trained is iteratively trained to obtain a target image processing model.
[0081] When iteratively training the image processing model to be trained, first obtain the first sample set; each first sample in the first sample set includes at least the video image to be trained and a first preset replacement image, and the first preset replacement image includes a replacement object for replacing the target object in the video image to be trained; then, based on the pre-constructed target occlusion image, perform occlusion on the first preset replacement image to obtain a corresponding second preset replacement image, and based on the second preset replacement image and the video image to be trained associated with the first preset replacement image, construct a second sample; therefore, for each first sample, a second sample associated therewith can be determined, and a second sample set is obtained based on the second sample. By constructing the first sample and the associated second sample, the scene of inter-frame moving object occlusion in the video can be simulated.
[0082] After obtaining the first sample set and the second sample set, based on the first sample set and the second sample set, the image processing model to be trained is iteratively trained to obtain a target image processing model; in one round of iterative training: first, select a first sample from the first sample set and a second sample associated with the first sample from the second sample set; then, input the first preset replacement image and the video image to be trained in the first sample into the image processing model to be trained, and through the image processing model to be trained, process the video image to be trained based on the first preset replacement image to obtain a first predicted image, and input the second preset replacement image and the video image to be trained in the second sample into the image processing model to be trained, and through the image processing model to be trained, process the video image to be trained based on the second preset replacement image to obtain a second predicted image; finally, adjust the parameters of the image processing model to be trained based on the inter-frame loss determined by the first predicted image and the second predicted image. By using the inter-frame loss to constrain the inter-frame stability of the video segment, the problem of inter-frame jitter during object replacement of each video image included in the video segment when there is a moving object occlusion in the video segment is prevented.
[0083] In summary, for the target image processing model obtained by training through the model training method provided in the embodiments of the present application, when replacing objects in each video image in a video clip, the inter-frame stability can be maintained, thereby ensuring the image quality of the processed video image.
[0084] The embodiments of the present application are designed based on computer vision technology (CV) and machine learning technology (ML) in artificial intelligence (AI).
[0085] Artificial intelligence is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0086] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0087] Computer vision technology is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition, monitoring, and measurement, and further performs graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0088] Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pretrained models are the latest development results of deep learning and integrate the above technologies.
[0089] The application scenarios set in this application are briefly described below. It should be noted that the following scenarios are only used to illustrate the embodiments of this application rather than to limit them. In specific implementations, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.
[0090] See Figure 1 , Figure 1 which is a schematic diagram of an application scenario provided by an embodiment of this application. This application scenario includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a communication network.
[0091] In an alternative embodiment, the communication network can be a wired network or a wireless network. Therefore, the terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication means. For example, the terminal device 110 can be indirectly connected to the server 120 through a wireless access point, or the terminal device 110 can be directly connected to the server 120 through the Internet. The present application does not limit this here.
[0092] Among them, the terminal device 110 includes but is not limited to devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.; various clients can be installed on the terminal device, and the client supports the video segment recording function.
[0093] The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The server 120 can also be a background server corresponding to the client installed in the terminal device 110.
[0094] In a possible application scenario, the terminal device 110 obtains the recorded video segment. When it is determined that the objects included in each video image in the recorded video segment need to be replaced, the objects to be replaced in the video segment are marked out, or the video images containing the objects to be replaced are selected from the video segment. Then, the video segment (or video image) and the replacement image are transmitted to the server 120, and the server 120 performs replacement processing using the target image processing model to obtain the processed target image. The target image processing model is pre-trained by the server 120 using the training method proposed in the embodiments of the present application.
[0095] It should be noted that the cooperation between the above terminal device 110 and the server 120 is only a possible implementation manner, and it can also be executed independently by the terminal device 110 or the server 120, which will not be elaborated here.
[0096] Figure 1 The above is only an example. In fact, the number of the terminal device 110 and the server 120 is not limited and is not specifically limited in the embodiments of the present application. In the embodiments of the present application, when the number of the servers 120 is multiple, the multiple servers 120 can form a blockchain, and the server 120 is a node on the blockchain.
[0097] See Figure 2 , Figure 2This is a specific application example diagram provided by an embodiment of the present application; from Figure 2 As can be seen, after the server 120 obtains the video image to be processed (the video image to be processed contains the face image of person A) and the target replacement image (the target replacement image contains the face image of person B), it inputs the image to be processed and the target replacement image into the target image processing model. The target image processing model outputs the processed target video image, which contains the facial expression, angle, and background of person A, and is quite similar to the face of person B, while maintaining high-definition picture quality.
[0098] To further illustrate the technical solution provided by the embodiment of the present application, the following takes the server's independent execution as an example and describes the training method of the image processing model provided by the exemplary embodiment of the present application in conjunction with the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0099] In the embodiment of the present application, the training process of the image processing model is a process of multi-round iterative training using the first sample set and the second sample set, which mainly includes a model design stage, a data preparation stage, and an iterative training stage. The following will introduce each stage separately.
[0100] I. Model Design Stage
[0101] See Figure 3 , Figure 3 This is a schematic structural diagram of an image processing model provided by an embodiment of the present application; the image processing model includes an encoder and a decoder, and the encoder includes at least one convolutional layer, which continuously halves the resolution and gradually increases the number of channels through convolutional calculation; the decoder includes at least one deconvolutional layer, which performs deconvolutional calculation to gradually increase the resolution and gradually reduce the number of channels.
[0102] As Figure 3 shown, the encoder includes 4 convolutional layers, and the decoder includes 4 deconvolutions; at this time, when performing replacement processing through the Figure 3 shown image processing model:
[0103] First, two images are concatenated as the input. Since the number of Red-Green-Blue (RGB) channels of each image is 3 and the resolution is 512*512, the input data for the input image processing model is 512*512*6; that is, the input data for the input encoder is 512*512*6, which is gradually encoded through 4 convolutional layers of the encoder. The encoding results from the first convolutional layer to the fourth convolutional layer are: 256*256*32, 128*128*64, 64*64*128, 32*32*256, respectively, to obtain an intermediate result in the latent space, called swap_features;
[0104] After that, in the latent space, the formula is used to fuse the swap features with the source image features (src_id_features) of the replacement image through Adaptive Instance Normalization (AdaIN) to obtain the corresponding fusion result. The size of the fusion result is the same as the data size of the last convolutional layer of the encoder, that is, 32*32*256; where x and y are the swap features and the source image features of the replacement image, respectively. It should be noted that the source image features of the replacement image are obtained through a trained feature extraction network;
[0105] Finally, the fusion result is input into the decoder, and is gradually decoded through 4 transposed convolutional layers of the decoder. The decoding results from the first transposed convolutional layer to the fourth transposed convolutional layer are: 64*64*128, 128*128*64, 256*256*32, 512*512*3, respectively, and finally the target image after replacement processing is obtained.
[0106] II. Data Preparation Stage
[0107] Data collection is of utmost importance in machine learning, and it can be said to be the most important link. The data preparation stage of this application embodiment mainly includes: the preparation process of the sample set and the construction process of the target occluded image.
[0108] 1. Preparation Process of Samples
[0109] The sample set includes a first sample set and a second sample set; the first sample set contains at least one first sample, and each first sample includes at least a first training image, a preset replacement image, and a first label image; the second sample set contains at least one second sample, and each second sample includes a second training image, a preset replacement image, and a second label image.
[0110] Since the second sample in this application embodiment is obtained based on the first sample, the acquisition method of the first sample will be introduced first.
[0111] In the embodiments of the present application, for various object replacement scenarios, an image triple in the object replacement scenario is obtained, and a first sample is determined based on the image triple; wherein, the image triple includes: an original image, a replacement image, and a generated image; the generated image is generated after replacing the original image with the replacement image.
[0112] Next, taking the replacement of a human face in a video clip as an example to illustrate the determination of the first sample. At this time, the image triple is: the original video image, the video replacement image, and the video generated image, and both the target object and the replacement object are human faces.
[0113] See Figure 4 , Figure 4 which is a schematic diagram for constructing a first sample provided by the embodiments of the present application, and specifically includes the following steps:
[0114] Step S400, obtain the original video image, the video replacement image, and the video generated image.
[0115] In a possible implementation manner, the original video image, the video replacement image, and the video generated image all include human face regions.
[0116] Step S401, perform face detection on the original video image, the video replacement image, and the video generated image respectively to obtain the corresponding human face regions.
[0117] Since in an image, a human face often occupies a relatively small position, it is necessary to perform face detection first to obtain the human face region.
[0118] Step S402, perform face registration on the original video image, the video replacement image, and the video generated image respectively within the corresponding human face regions to obtain the key points of the human face.
[0119] In a possible implementation manner, the key points include but are not limited to: eyes, nose, and mouth.
[0120] Step S403, crop the original video image, the video replacement image, and the video generated image respectively according to the key points identified within each human face region to obtain a first training image, a preset replacement image, and a first label image.
[0121] See Figure 5 , Figure 5 which is an example diagram of a first sample provided by the embodiments of the present application. The first sample includes: a first training image, a preset replacement image, and a first label image; it can be seen from Figure 5 that the first training image, the preset replacement image, and the first label image are of the same size, and most of the images are human face images with very little background image.
[0122] In this application, considering that some human faces in the image are skewed and the facial features information cannot be accurately collected, which may lead to inaccurate training results. Therefore, key point recognition is adopted, and the image is obtained according to the key point recognition result. The human face in the obtained image is upright, and the facial features information can be accurately recognized, ensuring the accuracy of the training samples during the training process and further ensuring the accuracy of the image processing model.
[0123] For ease of understanding, the embodiments of this application also provide an example diagram for constructing the first sample. Since the processing methods for the original video image, the video replacement image, and the video generated image are the same when performing image preprocessing to generate the corresponding first training image, the preset replacement image, and the first label image, only the construction of the first training image will be used as an example for display.
[0124] See Figure 6 , Figure 6 is a schematic diagram for constructing the first training image provided by the embodiments of this application; it can be seen from Figure 6 that for the obtained original video image: face detection needs to be performed first to obtain the face region, and then face registration is performed within the face region to obtain the key points of the face, with emphasis on the key points of the eyes and the corners of the mouth of the person. Finally, according to the face key points, the cropped face image is obtained, and the cropped face image is used as the first training image.
[0125] It should be noted that in different replacement scenarios, the target object and the replacement object are different. Therefore, the target object and the replacement object are determined according to the actual situation.
[0126] In the embodiments of this application, after obtaining the first sample, first, based on the pre-constructed target occlusion image, the first training image is occluded to obtain the corresponding second training image, and the first label image is occluded to obtain the corresponding second label image. Then, a second sample is constructed based on the second training image, the preset replacement image in the first sample, and the second label image.
[0127] See Figure 7 , Figure 7 is a schematic diagram for constructing the second sample based on the target occlusion image provided by the embodiments of this application; it can be seen from Figure 7 that:
[0128] First, a target occlusion image is randomly selected from the occlusion image library. The target occlusion image includes the occlusion object. It should be noted that each occlusion image in the occlusion image library is pre-constructed. For the specific construction process of the target occlusion image, please refer to the following, and it will not be elaborated here.
[0129] Then, perform the following operations on the first training image: Based on the mask image of the target occlusion image, perform occlusion processing on the first training image to obtain a masked training image, where the target area in the masked training image that matches the occlusion object is occluded; and based on the mask image of the target occlusion image, perform occlusion processing on the target occlusion image to obtain a masked occlusion image, where the other areas in the masked occlusion image except the target area that matches the occlusion object are occluded; superimpose the masked training image and the masked occlusion image to obtain a second training image;
[0130] That is: T2 = T1 * (1 - obj_mask) + obj_img * obj_mask, where T2 is the second training image, T1 is the first training image, obj_mask is the mask image, and obj_img is the target occlusion image;
[0131] Similarly, perform the following operations on the first label image: Based on the mask image of the target occlusion image, perform occlusion processing on the first label image to obtain a masked label image, where the target area in the masked label image that matches the occlusion object is occluded; and based on the mask image of the target occlusion image, perform occlusion processing on the target occlusion image to obtain a masked occlusion image, where the other areas in the masked occlusion image except the target area that matches the occlusion object are occluded; superimpose the masked label image and the masked occlusion image to obtain a second label image;
[0132] That is: GT2 = GT1 * (1 - obj_mask) + obj_img * obj_mask; where GT2 is the second label image, GT1 is the first label image, obj_mask is the mask image, and obj_img is the target occlusion image;
[0133] Among them, the mask image is used to determine the position area of the occlusion object in the target occlusion image;
[0134] Finally, construct a second sample based on the second training image, the second label image, and a preset replacement image.
[0135] In this application, by constructing the first sample and the associated second sample, the scenario of object occlusion during frame - to - frame movement in a video can be simulated.
[0136] 2. Construction process of the target occlusion image
[0137] Taking the scenario where a moving object in a video image occludes a face or the moving object moves on the face as an example, the method of obtaining the target occlusion image from the video image will be described.
[0138] Refer to Figure 8 , Figure 8 which is a flowchart of a method for constructing a target occlusion image provided by an embodiment of this application, including the following steps:
[0139] Step S800: Obtain the occluded image to be processed, where the occluded image to be processed includes at least an occluding object.
[0140] Step S801: Perform face detection on the occluded image to be processed to obtain the corresponding face region.
[0141] Since in an image, a face usually occupies a relatively small position, face detection needs to be performed first to obtain the face region.
[0142] Step S802: Perform face registration within the corresponding face region to obtain the key points of the face.
[0143] In a possible implementation, the key points include but are not limited to: eyes, nose, and mouth.
[0144] Step S803: Crop the occluded image to be processed according to the key points identified within the face region to obtain the cropped occluded image to be processed.
[0145] Step S804: Perform segmentation processing on the cropped occluded image to be processed through a pre-trained segmentation network to obtain a target occluded image and an associated mask image; the target occluded image includes the occluding object, and the mask image is used to determine the position of the occluding object in the target occluded image.
[0146] In a possible implementation, the segmentation network is applied to screen out the images with object occlusion and segment them to obtain a target occluded image (obj_img) and a mask image (obj_mask). Among them, the values of mask are 0 and 1, where 1 represents that there is an object at this position, and 0 represents no object. However, in the attached drawings, black filling represents that there is no object at this position, and white filling represents that there is an object at this position. These data serve as an occluded image library for subsequent data construction.
[0147] For ease of understanding, the embodiments of the present application provide a schematic diagram for constructing a target occluded image. Refer to Figure 9 , from Figure 9 it can be seen that:
[0148] The occluded image to be processed is an image of a star fragment occluding a face. At this time, face detection is first performed on the image to be processed to obtain the corresponding face region. Then, face registration is performed within the corresponding face region to obtain the key points of the face. Next, the occluded image to be processed is cropped according to the key points identified within the face region to obtain the cropped occluded image to be processed. Finally, the cropped occluded image to be processed is segmented through a pre-trained segmentation network to obtain a target occluded image only containing star fragments and an associated mask image.
[0149] It should be noted that, in order to ensure the accuracy of determining the second training image and the second label image based on the target occlusion image, when necessary, the image size of the target occlusion image is the same as that of the first training image and the first label image, that is, in the embodiments of the present application, the image sizes of all the involved images are the same.
[0150] In the embodiments of the present application, after the first sample for training and the associated second sample are prepared, the first sample and the associated second sample can be used to train the constructed model.
[0151] III. Iterative training stage
[0152] In the embodiments of the present application, the target image processing model is obtained by iteratively training the image processing model. During the model training process, a total of N rounds (such as 100) of iteration are performed on the full set of sample pairs. Among them, one round of iteration means that all the full set of sample pairs are trained once in the image processing model. In each round of iteration, due to the limited video memory resources of the training machine, the full set of sample pairs cannot be input into the model for training at one time. Therefore, all the sample pairs need to be trained in batches. Each batch of samples is generated by means such as random division, and each batch of samples is respectively input into the model for forward calculation, backward calculation, model parameter update and other training.
[0153] Before the first round of training, the parameters of the image processing model are initialized. Further, after setting the hyperparameters such as batch, epoch and learning rate respectively, the training starts, and finally the target image processing model is obtained.
[0154] See Figure 10 , Figure 10 which is a flowchart of a method for training an image processing model provided by the embodiments of the present application. This training method can be applied to a server or a terminal device, and includes the following steps:
[0155] Step S1000, obtain a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein, the preset replacement image includes: a replacement object for replacing the target object in the first training image.
[0156] Step S1001, for each first sample, based on the pre-constructed target occlusion image, perform occlusion processing on the first training image in the first sample to obtain a corresponding second training image, and based on the second training image and the preset replacement image in the first sample, construct a second sample associated with each first sample, and determine a second sample set based on the constructed second sample.
[0157] It should be noted that the specific implementation methods of steps S1000 and S1001 can refer to the implementation methods in the above data preparation stage, which will not be elaborated here.
[0158] Step S1002, select a first sample from the first sample set and a second sample associated with the first sample from the second sample set; wherein, the first sample and the second sample contain the same preset replacement image.
[0159] Step S1003, input the first sample and the second sample into the image processing model to be trained, and obtain a first predicted image and a second predicted image.
[0160] In a possible implementation manner, input the first training image and the preset replacement image of the first sample into the image processing model to be trained. In the image processing model to be trained, based on the replacement object in the preset replacement image, perform replacement processing on the target object in the first training image to obtain the first predicted image after replacement processing.
[0161] Similarly, input the second training image and the preset replacement image of the second sample into the image processing model to be trained. In the image processing model to be trained, based on the replacement object in the preset replacement image, perform replacement processing on the target object in the second training image to obtain the second predicted image after replacement processing.
[0162] It should be noted that the implementation method of obtaining the predicted image through the image processing model to be trained can refer to the model design stage, and the principle is similar, which will not be repeated here.
[0163] Step S1004, construct a target loss function based on the first predicted image and the second predicted image; wherein, the target loss function includes an inter-frame loss, and the inter-frame loss includes: an inter-frame feature loss and an inter-frame image loss.
[0164] In a possible implementation manner, determine the inter-frame feature loss in the following ways of steps A1 - A3:
[0165] Step A1, input the first predicted image and the second predicted image into the trained feature extraction network; the feature extraction network includes at least one network layer;
[0166] Step A2, in each network layer, determine the first image feature of the first predicted image and determine the second image feature of the second predicted image;
[0167] Step A3, determine the feature difference between the first image feature and the second image feature, and determine the inter-frame feature loss based on at least one feature difference.
[0168] Exemplarily, the feature extraction network is an AlexNet network. The AlexNet network includes 5 convolutional layers and 3 fully connected layers, and through this AlexNet network, features of an image at different layers can be extracted, including low-level features and high-level features. Low-level features can represent low-level features such as lines and colors, and high-level features can represent high-level features such as components. As Figure 11 shown, it is a schematic diagram of extracting features through a feature extraction network provided by an embodiment of the present application;
[0169] From Figure 11 it can be seen that through the AlexNet network, 4 groups of features can be determined for an image. Therefore, through the AlexNet network, 4 groups of features determined for the first predicted image are respectively: Image1-feature1, Image1-feature2, Image1-feature3, Image1-feature4; Similarly, through the AlexNet network, 4 groups of features determined for the second predicted image are respectively: Image2-feature1, Image2-feature2, Image2-feature3, Image2-feature4;
[0170] Since there are differences caused by the target occlusion image between the first predicted image and the second predicted image, if the feature differences between the first predicted image and the second predicted image are directly determined, the differences between them will surely be very large. Based on this difference value for training the model, the model training will be inaccurate. Therefore, to ensure accuracy, the embodiment of the present application only determines the feature differences between non-occluded regions, that is:
[0171] Image1-feature1, Image1-feature2, Image1-feature3, Image1-feature4 = AlexNet_feature(Image1 * (1 - obj_mask));
[0172] Image2-feature1, Image2-feature2, Image2-feature3, Image2-feature4 = AlexNet_feature(Image2 * (1 - obj_mask)).
[0173] Therefore, the inter-frame feature loss Occlu_frame_feature_loss = abs(Image1 - feature1 - Image2 - feature1) + abs(Image1 - feature2 - Image2 - feature2) + abs(Image1 - feature3 - Image2 - feature3) + abs(Image1 - feature4 - Image2 - feature4), where the abs function is a function for calculating the absolute value of data.
[0174] In a possible implementation, the inter-frame image loss is determined by the method of steps B1 - B3:
[0175] Step B1 determines the first non-occluded region in the first predicted image based on the mask image of the first predicted image and the target occluded image, that is, the first non-occluded region is Image1 * (1 - obj_mask);
[0176] Step B2 determines the second non-occluded region in the second predicted image based on the mask image of the second predicted image and the target occluded image, that is, the second non-occluded region is Image2 * (1 - obj_mask);
[0177] Step B3 determines the inter-frame image loss based on the regional image difference between the first non-occluded region and the second non-occluded region. The inter-frame image loss Occlu_frame_Image_loss = abs(Image1 * (1 - obj_mask) - Image2 * (1 - obj_mask)).
[0178] In a possible implementation, the inter-frame feature loss is the LPIPS loss; the inter-frame image loss is the L1 loss.
[0179] In the embodiments of the present application, in order to ensure that the predicted image is closer to the label image, and in the case of object movement and object occlusion, to ensure better processing effects and improve the quality of the processed image, the target loss function further includes: a prediction loss determined based on the difference between the first predicted image and the first label image, or a prediction loss determined based on the difference between the second predicted image and the second label image; the prediction loss includes: a prediction image loss and a prediction feature loss.
[0180] In a possible implementation, the prediction image loss Reconstruction_Image_loss = abs(Image1 - GT1), or Reconstruction_loss = abs(Image2 - GT2);
[0181] In a possible implementation, the predicted feature loss Reconstruction_feature_loss = abs(Image1 - feature1GT1 - feature1) + abs(Image1 - feature2GT1 - feature2) + abs(Image1 - feature3GT1 - feature3) + abs(Image1 - feature4GT1 - feature4), or the predicted feature loss Reconstruction_feature_loss = abs(Image2 - feature1 - GT2 - feature1) + abs(Image2 - feature2 - GT2 - feature2) + abs(Image2 - feature3 - GT2 - feature3) + abs(Image2 - feature4 - GT2 - feature4); where Image - feature1, Image - feature12, Image - feature1, Image - feature1 = AlexNet_feature(Image); GT_feature1, GT_feature2, GT_feature3, GT_feature4 = AlexNet_feature(GT).
[0182] In a possible implementation, the predicted feature loss is the LPIPS loss.
[0183] In the embodiments of the present application, in order to ensure that the predicted image is similar to the preset replacement image, the target loss function further includes an image generation loss; where the image generation loss is determined based on the similarity between the first predicted image and the preset replacement image, or based on the similarity between the second predicted image and the preset replacement image.
[0184] Exemplarily, the cosine similarity method is used to determine the image generation loss. The image generation loss: ID_loss = 1 - cosine_similarity(Image_id_features, source_id_features), where Image_id_features is the predicted image feature, source_id_features is the preset replacement image feature, and Image_id_features and source_id_features are object recognition features.
[0185] In the embodiments of the present application, in order to ensure that the predicted image is more realistic, the target loss function further includes an adversarial loss; that is, the image processing model to be trained in the embodiments of the present application is trained by using the Generative Adversarial Network (GAN).
[0186] In a possible implementation, the adversarial loss is determined in the following manner:
[0187] The first predicted image and the first labeled image in the first sample are input into the trained discriminative network to obtain the discrimination result of the first predicted image, and based on the discrimination result, the adversarial loss is determined; wherein, the discrimination result is used to characterize whether the first predicted image is a real image; or the second predicted image and the second labeled image in the second sample are input into the trained discriminative network to obtain the discrimination result of the second predicted image, and based on the discrimination result, the adversarial loss is determined; wherein, the discrimination result is used to characterize whether the second predicted image is a real image.
[0188] Exemplarily, the adversarial loss G_loss = log(1 - D((Image1)) or G_loss = log(1 - D((Image2)), where D((Image1)) and D((Image2)) are the discrimination results output by the discriminative network for the first predicted image and the second predicted image, respectively.
[0189] It should be noted that the discriminative network is trained based on the loss function D_loss = -logD(GT) - log(1 - D(Image)).
[0190] In summary, the target loss function in the embodiments of the present application is:
[0191] loss = a * Reconstruction_Image_loss + b * Reconstruction_feature_loss + c * ID_loss + d * G_loss + e * Occlu_frame_Image_loss + f * Occlu_frame_feature_loss; where a, b, c, d, e, and f are constants and can be obtained according to experience.
[0192] Step S1005: Based on the target loss function, adjust the parameters of the image processing model to be trained.
[0193] In some embodiments, only the first sample set may be constructed. During one iteration training process, first, a first sample is selected from the first sample set, and then, based on the pre-constructed target occlusion image, the first training image and the first label image are occluded to obtain the corresponding second training image and second label image. Then, the first training image, the second training image, the preset replacement image, the first label image, and the second label image are input into the image processing model to be trained. Through the image processing model to be trained, based on the replacement object in the preset replacement image, the target object in the first training image and the second training image are respectively replaced to obtain the corresponding first prediction image and second prediction image. And a frame-interpolation loss is constructed based on the first prediction image and the second prediction image, a generation loss is constructed based on the prediction image and the preset replacement image, a prediction loss and an adversarial loss are constructed based on the prediction image and the label image. Finally, a target loss function is constructed based on the frame-interpolation loss, the generation loss, the prediction loss, and the adversarial loss, and based on the target loss function, the parameters of the image processing model to be trained are adjusted.
[0194] For ease of understanding, an embodiment of the present application provides a training schematic diagram of an image processing model. Refer to Figure 12 , from Figure 12 it can be seen that:
[0195] First, the first sample is determined. Then, a target occlusion image is selected from the occlusion image library. Then, based on the pre-constructed target occlusion image, the first training image and the first label image are occluded to obtain the corresponding second training image and second label image. Then, the first training image, the second training image, the preset replacement image, the first label image, and the second label image are input into the image processing model to be trained;
[0196] In the image processing model to be trained, through the encoder and the decoder, based on the replacement object in the preset replacement image, the target object in the first training image and the second training image are respectively replaced to obtain the corresponding first prediction image and second prediction image;
[0197] A frame-interpolation loss is constructed based on the first prediction image and the second prediction image, a generation loss is constructed based on the prediction image and the preset replacement image, a prediction loss is constructed based on the prediction image and the label image, and the prediction image and the label image are input into the discriminant network, and an adversarial loss is constructed based on the output result of the discriminant network. Finally, a target loss function is constructed based on the frame-interpolation loss, the generation loss, the prediction loss, and the adversarial loss, and based on the target loss function, the parameters of the image processing model to be trained are adjusted.
[0198] In an embodiment of the present application, after obtaining the target image processing model, the video image to be processed and the target replacement image are input into the target image processing model, and the following operations are performed in the target image processing model: based on the replacement object in the target replacement image, the object to be replaced in the image to be processed is replaced to obtain the target video image after the replacement process. Exemplarily, as shown in Figure 2 which will not be elaborated here.
[0199] In the present application, when iteratively training the image processing model to be trained, first, a first sample set is obtained; each first sample in the first sample set at least includes a video image to be trained and a first preset replacement image, and the first preset replacement image includes a replacement object for replacing the target object in the video image to be trained; then, based on the pre-constructed target occlusion image, the first preset replacement image is occluded to obtain the corresponding second preset replacement image, and based on the second preset replacement image and the video image to be trained associated with the first preset replacement image, a second sample is constructed; thus, for each first sample, a second sample associated therewith can be determined, and a second sample set is obtained based on the second samples. By constructing the first samples and the associated second samples, the scenario of object occlusion during frame movement in a video can be simulated. After obtaining the first sample set and the second sample set, the image processing model to be trained is iteratively trained based on the first sample set and the second sample set to obtain the target image processing model; in one round of iterative training: first, a first sample is selected from the first sample set, and a second sample associated with the first sample is selected from the second sample set; then, the first preset replacement image and the video image to be trained in the first sample are input into the image processing model to be trained, and through the image processing model to be trained, based on the first preset replacement image, the video image to be trained is processed to obtain a first predicted image, and the second preset replacement image and the video image to be trained in the second sample are input into the image processing model to be trained, and through the image processing model to be trained, based on the second preset replacement image, the video image to be trained is processed to obtain a second predicted image; finally, based on the inter-frame loss determined by the first predicted image and the second predicted image, the parameters of the image processing model to be trained are adjusted. By using the inter-frame loss to constrain the inter-frame stability of the video segment, when replacing objects in each video image included in the video segment in the case of moving object occlusion in the video segment, the problem of inter-frame jitter can be prevented. Therefore, when the target image processing model obtained by training through the model training method proposed in the embodiment of the present application replaces objects in each video image in the video segment, the inter-frame stability can be maintained, thereby ensuring the image quality of the processed video image.
[0200] In addition, it should be noted that in the specific implementation of this application, when it comes to data related to users, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0201] Based on the same inventive concept, an embodiment of this application also provides a training device 1300 for an image processing model. Refer to Figure 13 The training device 1300 for the image processing model includes:
[0202] An acquisition unit 1301, configured to acquire a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein, the preset replacement image includes: a replacement object for replacing a target object in the first training image;
[0203] A processing unit 1302, configured to, for each first sample, perform occlusion processing on the first training image in the first sample based on a pre-constructed target occlusion image to obtain a corresponding second training image, and construct a second sample associated with each first sample based on the second training image and the preset replacement image in the first sample;
[0204] A training unit 1303, configured to perform iterative training on the image processing model to be trained based on the first sample set and the second sample set to obtain a target image processing model, wherein the second sample set includes the second samples associated with each first sample; in one round of iterative training, adjust the parameters of the image processing model to be trained based on the inter-frame loss determined by the first prediction image and the second prediction image; the first prediction image is: obtained by processing the corresponding first training image through the image processing model to be trained based on the preset replacement image in the selected first sample; the second prediction image is: obtained by processing the corresponding second training image through the image processing model to be trained based on the preset replacement image in the second sample associated with the selected first sample.
[0205] In a possible implementation manner, the target occlusion image includes an occlusion object; the processing unit 1302 is specifically configured to:
[0206] Perform occlusion processing on the first training image based on the mask image of the target occlusion image to obtain a masked training image; the target area in the masked training image that matches the occlusion object is occluded; the mask image is used to determine the position area of the occlusion object in the target occlusion image;
[0207] Perform occlusion processing on the target occlusion image based on the mask image of the target occlusion image to obtain a masked occlusion image; other areas in the masked occlusion image except the target area that matches the occlusion object are occluded;
[0208] Overlay the masked training image and the masked occlusion image to obtain a second training image.
[0209] In a possible implementation, the training unit 1303 is specifically configured to:
[0210] In one round of iterative training, perform the following operations:
[0211] Select a first sample from the first sample set and a second sample associated with the first sample from the second sample set; wherein, the first sample and the second sample contain the same preset replacement image;
[0212] Input the first sample and the second sample into the image processing model to be trained to obtain a first predicted image and a second predicted image;
[0213] Based on the first predicted image and the second predicted image, construct an objective loss function; wherein, the objective loss function includes an inter-frame loss;
[0214] Based on the objective loss function, adjust the parameters of the image processing model to be trained.
[0215] In a possible implementation, the inter-frame loss includes an inter-frame feature loss; and the inter-frame feature loss is determined by the following method:
[0216] Input the first predicted image and the second predicted image into a trained feature extraction network; wherein, the feature extraction network includes at least one network layer;
[0217] In each network layer, determine a first image feature of the first predicted image and a second image feature of the second predicted image;
[0218] Determine the feature difference between the first image feature and the second image feature, and determine the inter-frame feature loss based on at least one feature difference.
[0219] In a possible implementation, the inter-frame loss includes an inter-frame image loss; and the inter-frame image loss is determined by the following method:
[0220] Based on the first predicted image and the mask image of the target occlusion image, determine a first non-occluded region in the first predicted image;
[0221] Based on the second predicted image and the mask image of the target occlusion image, determine a second non-occluded region in the second predicted image;
[0222] Based on the regional image difference between the first non-occluded region and the second non-occluded region, determine the inter-frame image loss.
[0223] In a possible implementation, the first sample further includes a first labeled image, and the second sample further includes a second labeled image, where the second labeled image is obtained by performing an occlusion process on the first labeled image based on a target occlusion image;
[0224] The target loss function further includes: a prediction loss determined based on the difference between the first predicted image and the first labeled image, or a prediction loss determined based on the difference between the second predicted image and the second labeled image; the prediction loss includes: a predicted image loss and a predicted feature loss.
[0225] In a possible implementation, the target loss function further includes an image generation loss; and the image generation loss is determined in the following manner:
[0226] The image generation loss is determined based on the similarity between the first predicted image and a preset replacement image; or the image generation loss is determined based on the similarity between the second predicted image and a preset replacement image.
[0227] In a possible implementation, the target loss function further includes an adversarial loss; and the adversarial loss is determined in the following manner:
[0228] The first predicted image and the first labeled image in the first sample are input into a trained discriminative network to obtain a discrimination result of the first predicted image, and the adversarial loss is determined based on the discrimination result; where the discrimination result is used to characterize whether the first predicted image is a real image; or
[0229] The second predicted image and the second labeled image in the second sample are input into a trained discriminative network to obtain a discrimination result of the second predicted image, and the adversarial loss is determined based on the discrimination result; where the discrimination result is used to characterize whether the second predicted image is a real image.
[0230] In a possible implementation, the target occlusion image is obtained in the following manner:
[0231] Obtain an occlusion image to be processed, where the occlusion image to be processed includes at least an occlusion object;
[0232] The occlusion image to be processed is subjected to a segmentation process by a trained segmentation network to obtain a target occlusion image and an associated mask image; the target occlusion image includes the occlusion object, and the mask image is used to determine the position of the occlusion object in the target occlusion image.
[0233] In a possible implementation, the video image to be trained and the first preset replacement image are obtained after image preprocessing; the video image to be trained contains a target object, and the first preset replacement image contains a replacement object.
[0234] In a possible implementation, after the training unit 1303 obtains the target image processing model, it is further configured to:
[0235] Based on the replacement object in the target replacement image through the target image processing model, perform replacement processing on the object to be replaced in the video image to be processed, and obtain the target video image.
[0236] It should be noted that although several units (or modules) of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units (or modules) described above can be embodied in one unit (or module). Conversely, the features and functions of one unit (or module) described above can be further divided and embodied by multiple units (or modules). Of course, when implementing the present application, the functions of each unit (or module) can also be implemented in the same or multiple software or hardware.
[0237] In the embodiments of the present application, the term unit (or module) refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0238] After introducing the training method and device of the image processing model in the exemplary embodiments of the present application, the following introduces another exemplary embodiment of the present application, the computing device.
[0239] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0240] In a possible implementation, the computing device provided in the embodiments of the present application may at least include a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any step in the training method of the image processing model in various exemplary embodiments of the present application.
[0241] In this embodiment, the structure of the computing device can be as Figure 14As shown, it includes a memory 1401, a communication module 1403, and one or more processors 1402.
[0242] The memory 1401 is used to store computer programs executed by the processor 1402. The memory 1401 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.
[0243] The memory 1401 can be a volatile memory, such as a random-access memory (RAM); the memory 1401 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1401 is any other medium that can be used to carry or store a desired computer program in the form of instruction or data structure and can be accessed by a computer, but not limited thereto. The memory 1401 can be a combination of the above memories.
[0244] The processor 1402 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1402 is used to implement the above-mentioned training method of the image processing model when calling the computer program stored in the memory 1401.
[0245] The communication module 1403 is used to communicate with terminal devices and other servers.
[0246] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1401, communication module 1403, and processor 1402 is not limited. In the embodiments of the present application Figure 14 it is described that the memory 1401 and the processor 1402 are connected through a bus 1404, and the bus 1404 is described in thick lines in Figure 14 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1404 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 14 it is only described by a thick line in
[0247] The computer storage medium is stored in the memory 1401. The computer-executable instructions are stored in the computer storage medium, and the computer-executable instructions are used to implement the training method of the image processing model in the embodiments of the present application. The processor 1402 is used to execute the above-mentioned training method of the image processing model.
[0248] In some possible implementation manners, aspects of the training method of the image processing model provided by the present application may also be implemented in the form of a program product, which includes a computer program. When the program product runs on a computing device, the computer program is used to cause the computing device to execute the steps in the training method of the image processing model according to various exemplary embodiments of the present application described above in this specification.
[0249] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0250] The program product of the embodiments of the present application may adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and may run on a computing device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with a command execution system, apparatus, or device.
[0251] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, and the readable medium may send, propagate, or transmit a program for use by or in combination with a command execution system, apparatus, or device.
[0252] The computer program included on the readable medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0253] The computer program for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0254] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable computer programs.
[0255] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable devices generate means for realizing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0256] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0257] These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the specified functions in one block or a plurality of blocks.
[0258] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0259] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A training method for an image processing model, characterized in that: The method comprises: Acquire a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein the preset replacement image includes: a replacement object for replacing a target object in the first training image; For each of the first samples, based on a pre-constructed target occlusion image, perform occlusion processing on the first training image in the first sample to obtain a corresponding second training image, and construct a second sample associated with each of the first samples based on the second training image and a preset replacement image in the first sample; Based on the first sample set and the second sample set, the image processing model to be trained is iteratively trained to obtain a target image processing model, wherein the second sample set includes a second sample associated with each of the first samples; in one round of iterative training, the parameters of the image processing model to be trained are adjusted based on the first predicted image and the inter-frame loss determined by the second predicted image; the first predicted image is obtained by processing the corresponding first training image through the image processing model to be trained based on the preset replacement image in the selected first sample; the second predicted image is obtained by processing the corresponding second training image through the image processing model to be trained based on the preset replacement image in the second sample associated with the selected first sample.
2. The method according to claim 1, characterized in that The target occlusion image includes an occlusion object; The step of performing occlusion processing on the first training image in the first sample based on the pre-constructed target occlusion image to obtain a corresponding second training image includes: Based on the mask image of the target occlusion image, the first training image is subjected to occlusion processing to obtain a shielded training image; the target area matching the occlusion object in the shielded training image is occluded; the mask image is used to determine the position area of the occlusion object in the target occlusion image; Based on the mask image of the target occlusion image, the target occlusion image is subjected to occlusion processing to obtain a shielded occlusion image; in the shielded occlusion image, other areas except the target area matching the occlusion object are occluded; The shielding training image and the shielding occlusion image are superimposed to obtain the second training image.
3. The method according to claim 1, characterized in that The iterative training of the image processing model to be trained based on the first sample set and the second sample set to obtain a target image processing model includes: In one round of iterative training, the following operations are performed: Selecting a first sample from the first sample set, and selecting a second sample associated with the first sample from the second sample set; wherein the first sample and the second sample contain the same preset replacement image; Inputting the first sample and the second sample into the image processing model to be trained to obtain the first predicted image and the second predicted image; Based on the first predicted image and the second predicted image, construct a target loss function; wherein the target loss function includes inter-frame loss; Based on the target loss function, parameters of the image processing model to be trained are adjusted.
4. The method according to claim 1 or 3, characterized in that The inter-frame loss includes inter-frame feature loss; And the inter-frame feature loss is determined by the following method: Inputting the first predicted image and the second predicted image into a trained feature extraction network; wherein the feature extraction network includes at least one network layer; In each network layer, determining a first image feature of the first predicted image, and determining a second image feature of the second predicted image; A feature difference between the first image feature and the second image feature is determined, and based on at least one feature difference, the inter-frame feature loss is determined.
5. The method according to claim 1 or 3, characterized in that: The inter-frame loss includes an inter-frame image loss; and the inter-frame image loss is determined by: Determining a first non-occluded area in the first predicted image based on the first predicted image and a mask image of the target occluded image; Determining a second non-occluded area in the second predicted image based on the second predicted image and the mask image of the target occluded image; The inter-frame image loss is determined based on a regional image difference between the first non-occluded area and the second non-occluded area.
6. The method according to claim 3, characterized in that The first sample further includes a first label image, and the second sample further includes a second label image, wherein the second label image is obtained by performing the occlusion processing on the first label image based on the target occlusion image; The target loss function also includes: a prediction loss determined based on a difference between the first predicted image and the first label image, or a prediction loss determined based on a difference between the second predicted image and the second label image; the prediction loss includes: a prediction image loss and a prediction feature loss.
7. The method according to claim 3, characterized in that The objective loss function also includes image generation loss; and the image generation loss is determined by the following method: The image generation loss is determined based on the similarity between the first predicted image and the preset replacement image; or the image generation loss is determined based on the similarity between the second predicted image and the preset replacement image.
8. The method according to claim 6 or 7, characterized in that The objective loss function also includes adversarial loss, and the adversarial loss is determined by the following method: Inputting the first predicted image and the first label image in the first sample into a trained discriminant network to obtain a discriminant result of the first predicted image, and determining the adversarial loss based on the discriminant result; wherein the discriminant result is used to characterize whether the first predicted image is a real image; or The second predicted image and the second label image in the second sample are input into a trained discriminant network to obtain a discrimination result of the second predicted image, and the adversarial loss is determined based on the discrimination result; wherein the discrimination result is used to characterize whether the second predicted image is a real image.
9. The method according to any one of claims 1 to 3, characterized in that: The target occlusion image is obtained by: Acquire an occlusion image to be processed, where the occlusion image to be processed at least includes an occlusion object; The occluded image to be processed is segmented by a trained segmentation network to obtain the target occluded image and an associated mask image; the target occluded image includes the occluded object, and the mask image is used to determine the position of the occluded object in the target occluded image.
10. The method according to any one of claims 1 to 3, characterized in that: The first training image and the preset replacement image are obtained after image preprocessing; the first training image contains the target object, and the preset replacement image contains the replacement object.
11. The method according to claim 1, characterized in that After obtaining the target image processing model, the method further includes: By using the target image processing model, based on the replacement object in the target replacement image, the object to be replaced in the processed video image is replaced to obtain the target video image.
12. A training device for an image processing model, characterized in that: The device comprises: An acquisition unit is used to acquire a first sample set; each first sample in the first sample set includes at least a first training image and a preset replacement image; wherein the preset replacement image includes: a replacement object for replacing a target object in the first training image; a processing unit, configured to, for each of the first samples, perform occlusion processing on the first training image in the first sample based on a pre-constructed target occlusion image to obtain a corresponding second training image, and construct a second sample associated with each of the first samples based on the second training image and a preset replacement image in the first sample; A training unit is used to iteratively train the image processing model to be trained based on the first sample set and the second sample set to obtain a target image processing model, wherein the second sample set includes a second sample associated with each of the first samples; in one round of iterative training, the parameters of the image processing model to be trained are adjusted based on the first predicted image and the inter-frame loss determined by the second predicted image; the first predicted image is obtained by processing the corresponding first training image through the image processing model to be trained based on the preset replacement image in the selected first sample; the second predicted image is obtained by processing the corresponding second training image through the image processing model to be trained based on the preset replacement image in the second sample associated with the selected first sample.
13. A computing device, characterized in that: The computing device comprises: a processor and a memory, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.