Model training data generation method and device, equipment, medium and product
By improving the consistency between the face stylized data and the original image, the problem of insufficient consistency in the existing technology caused by model training failure is solved, and effective training and terminal-side deployment of lightweight face stylized models are realized.
Patent Information
- Application Number
- CN202510306387.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The face stylized data generated in the prior art have low consistency with the original image, resulting in the failure of training of the face stylized model.
The original image data of the subject and the visual effect data required by the visual effect model are used to train the preset large parameter quantity generation model, and the trained large parameter quantity generation model is obtained. Then, the original image data is processed based on the model, the target subject visual effect data is generated, and it is used to generate the training data of the visual effect model.
The consistency between the face stylized data and the original image is improved, making it possible to lightweight stylized training based on the diffusion model, avoiding the risk of model training failure, and supporting the deployment of lightweight visual effects models on the terminal side.
Smart Images

Figure CN120219882A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, medium, and product for generating model training data. Background Art
[0002] Visual effects play an important role in fields such as movies, advertisements, and games. For example, portrait beautification, landscape rendering, face stylization, and so on.
[0003] Taking face stylization as an example, in order to enable end users to also experience rich face stylization effects, a lightweight face stylization model can be run in real time on the terminal side to achieve this. In related technologies, the lightweight face stylization model deployed on the terminal side is trained by using a large amount of face stylization data generated by a diffusion model (Diffusion Models, DM) as training data.
[0004] However, the face stylization data generated in related technologies has the technical problem of low consistency between the face stylization data and the original Figure 1 image, which easily leads to the failure of face stylization model training. Summary of the Invention
[0005] Based on this, in order to solve the above technical problems, it is necessary to provide a method, apparatus, device, medium, and product for generating model training data, which can improve the consistency between face stylization data and the original image, and make it possible to perform lightweight stylization training based on data generated by a diffusion model.
[0006] In a first aspect, an embodiment of the present application provides a method for generating model training data, including:
[0007] Using the main original image data and main visual effect image data required by the visual effect model to train a preset large-parameter generation model, and obtaining a trained large-parameter generation model; the main visual effect image data is obtained by performing visual effect processing on the main original image data through a diffusion model;
[0008] Processing the main original image data according to the trained large-parameter generation model to obtain target main visual effect image data;
[0009] Generating training data for the visual effect model according to the target main visual effect image data and the main original image data.
[0010] In one embodiment, using the main original image data and main visual effect image data required by the visual effect model to train a preset large-parameter generation model, and obtaining a trained large-parameter generation model, includes:
[0011] Input the original main image data into a large-parameter generation model to obtain a predicted main body visual effect image output by the large-parameter generation model;
[0012] Determine the trained large-parameter generation model according to the predicted main body visual effect image and the main body visual effect image data.
[0013] In one embodiment, determining the trained large-parameter generation model according to the predicted main body visual effect image and the main body visual effect image data includes:
[0014] Determine the prediction loss of the large-parameter generation model according to the predicted main body visual effect image, the main body visual effect image data, and a preset loss function;
[0015] Based on the prediction loss, adjust the model parameters of the large-parameter generation model until the training is completed to obtain the trained large-parameter generation model.
[0016] In one embodiment, generating the training data of the visual effect model according to the target main body visual effect image data and the original main image data includes:
[0017] Perform main body consistency screening and background consistency processing on the target main body visual effect image data to obtain processed main body visual effect image data;
[0018] Generate the training data of the visual effect model according to the processed main body visual effect image data and the original main image data.
[0019] In one embodiment, performing main body consistency screening and background consistency processing on the target main body visual effect image data to obtain processed main body visual effect image data includes:
[0020] Remove the target main body visual effect images in the target main body visual effect image data that have a large difference from the original main body according to the original main image data and the target main body visual effect image data to obtain a main body consistency visual effect image set;
[0021] Replace the visual effect background in the main body consistency visual effect image set with the original image background according to the original main image data and the target main body visual effect image data to obtain the processed main body visual effect image data.
[0022] In one embodiment, removing the target main body visual effect images in the target main body visual effect image data that have a large difference from the original main body according to the original main image data and the target main body visual effect image data to obtain a main body consistency visual effect image set includes:
[0023] Perform main body segmentation processing on the original main image data to obtain an original main body mask set; and perform main body segmentation processing on the target main body visual effect image data to obtain a visual effect main body mask set;
[0024] Based on the original main body mask set and the visual effect main body mask set, perform image removal processing on the target main body visual effect map data to obtain a main body consistency visual effect map set.
[0025] In one embodiment, based on the original main body mask set and the visual effect main body mask set, performing image removal processing on the target main body visual effect map data to obtain a main body consistency visual effect map set includes:
[0026] Based on the original main body mask set and the visual effect main body mask set, determine the main body consistency decision data for each target main body visual effect map in the target main body visual effect map data;
[0027] Based on each main body consistency decision data, perform image removal processing on the target main body visual effect map data to obtain a main body consistency visual effect map set.
[0028] In one embodiment, based on the original main body mask set and the visual effect main body mask set, determining the main body consistency decision data for each target main body visual effect map in the target main body visual effect map data includes:
[0029] For any target main body visual effect map, based on the original main body mask set and the visual effect main body mask set, obtain the corresponding original main body mask and the corresponding visual effect main body mask of the target main body visual effect map;
[0030] Determine the overlapping area of the original main body mask and the visual effect main body mask;
[0031] Determine the ratio of the overlapping area to the original main body mask as the main body consistency decision data of the target main body visual effect map.
[0032] In one embodiment, based on each main body consistency decision data, performing image removal processing on the target main body visual effect map data to obtain a main body consistency visual effect map set includes:
[0033] Obtain a preset main body consistency threshold;
[0034] Remove the target main body visual effect maps whose main body consistency decision data is less than the main body consistency threshold;
[0035] Determine the remaining target main body visual effect maps in the target main body visual effect map data as the main body consistency visual effect map set.
[0036] In one embodiment, based on the main body original image data and the target main body visual effect map data, replace the visual effect background in the main body consistency visual effect map set with the original image background to obtain the processed main body visual effect map data, including:
[0037] Perform subject segmentation on the original subject image data to obtain the original subject mask set; and, perform subject segmentation on the target subject visual effect image data to obtain the visual effect subject mask set;
[0038] According to the original subject mask set and the visual effect subject mask set, determine the background consistency decision data for each subject consistency visual effect map in the subject consistency visual effect map set;
[0039] According to the background consistency decision data, perform visual effect background replacement on the subject consistency visual effect map set to obtain the processed subject visual effect image data.
[0040] In one embodiment, determining the background consistency decision data for each subject consistency visual effect map in the subject consistency visual effect map set according to the original subject mask set and the visual effect subject mask set includes:
[0041] For any subject consistency visual effect map, obtain the corresponding original subject mask and the corresponding visual effect subject mask according to the original subject mask set and the visual effect subject mask set;
[0042] Determine the union mask of the original subject mask and the visual effect subject mask, and perform dilation operation and Gaussian blur processing on the union mask to obtain the processed union mask;
[0043] Determine the processed union mask as the background consistency decision data for the subject consistency visual effect map.
[0044] In one embodiment, performing visual effect background replacement on the subject consistency visual effect map set according to the background consistency decision data to obtain the processed subject visual effect image data includes:
[0045] Obtain the preset background replacement mathematical model and the corresponding subject consistency original map set of the subject consistency visual effect map set;
[0046] According to the background consistency decision data, the background replacement mathematical model and the subject consistency original map set, perform visual effect background replacement on the subject consistency visual effect map set to obtain the processed subject visual effect image data.
[0047] In one embodiment, the method further includes:
[0048] Obtain multiple historical original subject images;
[0049] Input each historical original subject image into the diffusion model to obtain multiple subject visual effect images output by the diffusion model;
[0050] Determine the original subject images of each historical subject as the original subject image data, and determine the visual effect images of each subject as the visual effect subject image data.
[0051] In one embodiment, before training a preset large-parameter generation model using the original subject image data and the visual effect subject image data to obtain a trained large-parameter generation model, the method further includes:
[0052] Obtain the original subject mask set of the original subject image data and the visual effect subject mask set of the visual effect subject image data;
[0053] Perform subject consistency screening on the original subject image data and the visual effect subject image data according to the subject mask and the visual effect subject mask to obtain the screened original subject image data and the screened visual effect subject image data.
[0054] In one embodiment, the method further includes:
[0055] Obtain target original subject image data according to the target visual effect subject image data and the original subject image data;
[0056] Input the target original subject image data into a preset small-parameter generation model to obtain predicted visual effect subject image data output by the small-parameter generation model;
[0057] Adjust the model parameters of the small-parameter generation model according to the predicted visual effect subject image data, the target visual effect subject image data, and a preset loss function until the training is completed to obtain a visual effect model.
[0058] In one embodiment, the method further includes:
[0059] Obtain an original subject image to be processed;
[0060] Input the original subject image into the visual effect model to obtain the subject visual effect image of the original subject image.
[0061] In a second aspect, an embodiment of the present application further provides a device for generating model training data, including:
[0062] A model training module, configured to train a preset large-parameter generation model using the original subject image data and the visual effect subject image data required by the visual effect model to obtain a trained large-parameter generation model; the visual effect subject image data is obtained by performing visual effect processing on the original subject image data through a diffusion model;
[0063] A data processing module, configured to process the original subject image data according to the trained large-parameter generation model to obtain target visual effect subject image data;
[0064] A training data generation module, configured to generate training data for the visual effect model according to the target subject visual effect map data and the original subject map data.
[0065] In a third aspect, an embodiment of the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in any one of the embodiments in the first aspect are implemented.
[0066] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the embodiments in the first aspect are implemented.
[0067] In a fifth aspect, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the embodiments in the first aspect are implemented.
[0068] The method, device, equipment, medium and product for generating model training data provided by the embodiments of the present application use the original subject map data and the target subject visual effect map data required by the visual effect model to train a preset large-parameter generation model, and obtain a trained large-parameter generation model. Then, the original subject map data is processed according to the trained large-parameter generation model to obtain the target subject visual effect map data. After that, the training data of the visual effect model is generated according to the target subject visual effect map data and the original subject map data, where the target subject visual effect map data is obtained by performing visual effect processing on the original subject map data through a diffusion model. In this method, the original subject map data and the target subject visual effect map data obtained by the diffusion model are used to train a preset large-parameter generation model. The large-parameter generation model has a powerful learning ability due to its large number of parameters, so that the large-parameter generation model can learn more details and features of the original subject map data. Therefore, when performing visual effect processing, it can more accurately copy the details of the original subject map data, and then the trained large-parameter generation model can generate a visual effect map with higher consistency with the original subject map data. Based on this, when the original subject map data is processed by the trained large-parameter generation model, the obtained target subject visual effect map data will have a high consistency with the original subject map data. In this way, using the target subject visual effect map data for model training is not likely to cause model training failure; and in order to train a lightweight visual effect model that can run on the terminal side, the target subject visual effect map data needs to be trained in a small-parameter model. Since the target subject visual effect map data has a high consistency with the original subject map data, it will not interfere with the training of the small-parameter model, making it possible to deploy a lightweight visual effect model on the terminal side, and thus enabling terminal users to experience richer visual effects. Brief Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 It is the internal structure diagram of a computer device in an embodiment;
[0071] Figure 2 It is the application environment diagram of the method for generating model training data in an embodiment;
[0072] Figure 3 It is the internal structure diagram of a computer device in another embodiment;
[0073] Figure 4 It is the flow schematic diagram of the method for generating model training data in an embodiment;
[0074] Figure 5 It is the flow schematic diagram of training a large-parameter generation model in an embodiment;
[0075] Figure 6 It is the flow schematic diagram of determining the trained large-parameter generation model in an embodiment;
[0076] Figure 7 It is the flow schematic diagram of generating training data for a visual effect model in an embodiment;
[0077] Figure 8 It is the flow schematic diagram of determining the processed main body visual effect map data in an embodiment;
[0078] Figure 9 It is the flow schematic diagram of determining the main body consistency visual effect map set in an embodiment;
[0079] Figure 10 It is the flow schematic diagram of determining the main body consistency visual effect map set in another embodiment;
[0080] Figure 11 It is the flow schematic diagram of determining the main body consistency decision data in an embodiment;
[0081] Figure 12 It is the flow schematic diagram of determining the main body consistency visual effect map set in another embodiment;
[0082] Figure 13Schematic diagram of the process for determining the visual special effect map data of the processed subject in another embodiment;
[0083] Figure 14 Schematic diagram of the process for determining the background consistency decision data in one embodiment;
[0084] Figure 15 Schematic diagram of the process for determining the visual special effect map data of the processed subject in another embodiment;
[0085] Figure 16 Schematic diagram of the process for determining the original subject image data and the visual special effect map data of the subject in one embodiment;
[0086] Figure 17 Schematic diagram of the process for subject consistency screening of the original subject image data and the visual special effect map data of the subject in one embodiment;
[0087] Figure 18 Schematic diagram of the process for constructing a visual special effect model in one embodiment;
[0088] Figure 19 Schematic diagram of the process for determining the visual special effect image of the subject in one embodiment;
[0089] Figure 20 Schematic diagram of the process for the generation method of model training data in another embodiment;
[0090] Figure 21 Schematic diagram of the process for determining a lightweight face stylization model in one embodiment;
[0091] Figure 22 Structural block diagram of the device for generating model training data in one embodiment. Detailed implementation manners
[0092] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0093] The technical background of the present application will be described below first.
[0094] Visual special effects play an important role in fields such as movies, TV dramas, advertisements, games, etc. For example, portrait beauty / filters, landscape rendering, face stylization, and so on. Among them, face stylization is a common visual special effect, and this technology is applicable not only to static images but also to face transformation in dynamic videos to achieve real-time face animation effects.
[0095] Taking face stylization as an example, in order to enable smart terminal users to also experience rich face stylization effects, a lightweight face stylization model can be run in real time on the terminal side to achieve this. The lightweight face stylization model requires a large amount of face stylization data during training. In related technologies, the lightweight face stylization model deployed on the terminal side is trained using a large amount of face stylization data generated by a diffusion model with stronger stylization capabilities as training data. However, the consistency between the face stylization data generated by the diffusion model and the original image is relatively low. Directly training the lightweight face stylization model on this data will lead to difficult convergence and learning failure, that is, it is easy to cause the training of the face stylization model to fail.
[0096] Based on this, the embodiment of the present application provides a method for generating model training data. Using the main original image data and the main body visual effect map data obtained by the diffusion model, a preset large-parameter generation model is trained. Due to the large number of parameters, the large-parameter generation model has a strong learning ability, enabling the large-parameter generation model to learn more details and features of the main original image data. Thus, when performing visual effect processing, it can more accurately copy the details of the main original image data. Furthermore, the trained large-parameter generation model can generate a visual effect map with higher consistency with the main original image data. Based on this, by processing the main original image data with the trained large-parameter generation model, the obtained target main body visual effect map data will have a high consistency with the main original image data. In this way, using the target main body visual effect map data for model training is not likely to cause model training failure; and in order to be able to train a lightweight visual effect model that runs on the terminal side, the target main body visual effect map data needs to be trained in a small-parameter model. Due to the high consistency between the target main body visual effect map data and the main original image data, it will not interfere with the training of the small-parameter model, making it possible to deploy a lightweight visual effect model on the terminal side, and thus enabling terminal users to experience richer visual effects.
[0097] It should be noted that the beneficial effects or the technical problems solved by the embodiments of the present application are not limited to this one, and there may be other implicit or related problems. For specific details, please refer to the descriptions in the following embodiments.
[0098] Next, the application environment of the method for generating model training data provided by the embodiments of the present application will be described.
[0099] The method for generating model training data provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, this method for generating model training data is applied in a Figure 1 shown computer device. This computer device can be a server. For its internal structure, please refer to Figure 1The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the generated data of the model training data. The network interface of the computer device is used to exchange information with external devices. When the computer program is executed by the processor, it implements a method for generating model training data.
[0100] In addition to the above scenarios, the method for generating model training data provided in the embodiments of the present application can also be applied to, for example Figure 2 the application environment shown. Among them, the terminal 102 communicates with the server 104 via a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. In combination with the embodiments of the present application, the terminal 102 can be used to send the original subject image data required by the visual effect model to the server 104, so that the server 104 processes the original subject image data into the subject visual effect image data through the diffusion model. Furthermore, the server 104 uses the original subject image data and the subject visual effect image data to train a preset large-parameter generation model, and uses the trained large-parameter generation model to process the original subject image data into the target subject visual effect image data, so as to generate the training data of the visual effect model based on the target subject visual effect image data and the original subject image data.
[0101] In addition, the method for generating model training data provided in the embodiments of the present application can also be applied to, for example Figure 3 the application environment shown. Among them, the method for generating the model training data is applied in Figure 3 the computer device shown. The computer device can be a terminal, and its internal structure is shown in Figure 3The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a method for generating model training data. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0102] Those skilled in the art can understand that Figure 1 and Figure 3 the structure shown in
[0103] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0104] In an exemplary embodiment, as Figure 4 shown, a method for generating model training data is provided. Taking the method applied to Figure 1 the server 104 in
[0105] S201. Use the original subject image data and the subject visual effect image data required by the visual effect model to train a preset large-parameter generation model to obtain a trained large-parameter generation model. The subject visual effect image data is obtained by performing visual effect processing on the original subject image data through a diffusion model.
[0106] Among them, the visual effect model refers to a lightweight model that can run on the intelligent terminal side for visual effect processing. Taking face stylization as an example of visual effect, the visual effect model refers to a lightweight face stylization model. Taking portrait beautification as an example of visual effect, the visual effect model refers to a lightweight portrait beautification model. Taking landscape rendering as an example of visual effect, the visual effect model refers to a lightweight landscape rendering model.
[0107] The original subject image data refers to the original image data without any processing. The subject refers to the main object in the image, and the subject can be a person or an object. For example, in portrait photography, the person is the subject. In scene photography, a main scene or person can be selected as the subject. In landscape photography, a representative object can be selected as the subject.
[0108] The subject visual effect image data refers to the image data obtained by performing visual effect processing on the original subject image data. Taking face stylization as an example of visual effect, the original subject image data is the original face image, and the subject visual effect image data is the face stylized image obtained by performing stylization processing on the original face image. In the embodiments of the present application, the subject visual effect image data is obtained by performing visual effect processing on the original subject image data through a diffusion model. Taking face stylization as an example of visual effect, the original subject image data is the original face image, and the subject visual effect image data is the face stylized image. The diffusion model can perform stylization processing on the original face image to obtain the face stylized image.
[0109] The diffusion model has powerful visual effect processing capabilities. However, the multi-step denoising process of the diffusion model will cause inconsistencies in details between the visually processed image and the original image. Taking the face as an example, parts such as the face shape, hair position, eye orientation, and mouth movement may be inconsistent with the original image. And the small-parameter model has fewer parameters and weaker learning ability. This inconsistency is likely to cause great interference to the training of the small-parameter model, resulting in difficulty in convergence. Therefore, in the embodiments of the present application, the large-parameter generation model learns the visual effect processing ability, so that the large-parameter generation model can capture more details and features in the original image based on its own powerful learning ability. Thus, when generating a visual effect image, it can more accurately replicate the style and details of the original image, improving the consistency between the visual effect image and the original image. Based on this, compared with the diffusion model, the large-parameter generation model can generate a visual effect image with higher Figure 1 consistency with the original, so that it will be the same as the original Figure 1Visual effect maps with higher consistency are used for small-parameter models and will not interfere with the training of small-parameter models, that is, they will not cause the training of the visual effect model to fail, thus making it possible to perform lightweight visual effect processing from the data generated by the diffusion model to the terminal side.
[0110] Among them, the large-parameter generation model refers to a machine learning model with a huge number of parameters, usually constructed by a deep neural network, having billions or even hundreds of billions of parameters. The purpose of this design is mainly to improve the expression ability and prediction performance of the model and can handle more complex tasks and data. In the embodiments of the present application, the large-parameter generation model can be a large-parameter Generative Adversarial Networks (GAN), and the number of parameters can be more than 300M.
[0111] The training of the large-parameter generation model requires the original subject image data and the subject visual effect map data generated by the diffusion model as training data. Therefore, it is necessary to obtain the original subject image data and the subject visual effect map data.
[0112] In one embodiment, the way to obtain the original subject image data can be to obtain it from a database. For example, there are multiple historical original subject image data stored in the database. All the historical original subject image data can be obtained from the database as the original subject image data required by the visual effect model, or a part of the historical original subject image data can be selected from the database as the original subject image data required by the visual effect model. This is not limited in the embodiments of the present application.
[0113] In another embodiment, the way to obtain the original subject image data can be to obtain it through a data acquisition interface. For example, when it is necessary to obtain the original subject image data required by the visual effect model, a data acquisition interface can be displayed to the user, and then the user uploads the original subject image data to the data acquisition interface. In this way, the server can obtain the original subject image data.
[0114] After obtaining the original subject image data, the original subject image data is input into the diffusion model, and the diffusion model performs visual effect processing on the original subject image data to obtain the subject visual effect map data.
[0115] Based on the obtained original subject image data and the subject visual effect map data, a preset large-parameter generation model is trained to obtain a trained large-parameter generation model. For example, the original subject image data can be input into the preset large-parameter generation model to obtain a prediction result, and then a loss calculation is performed based on the prediction result and the subject visual effect map data, and the model parameters of the large-parameter generation model are continuously adjusted based on the calculated loss to obtain a trained large-parameter generation model.
[0116] S202. Process the original subject image data according to the trained large-parameter generation model to obtain the target subject visual effect image data.
[0117] Using the original subject image data and the target subject visual effect image data to train the large-parameter generation model, it can learn the visual effect processing ability and obtain a visual effect image with higher consistency with the original. Figure 1 Based on this, the trained large-parameter generation model can be used to perform visual effect processing on the original subject image data to obtain visual effect image data with higher consistency with the original subject image data, which can be used for the training of the visual effect model.
[0118] In the embodiment of the present application, the original subject image data can be input into the trained large-parameter generation model to obtain the target subject visual effect image data, and the target subject visual effect image data has a high consistency with the original subject image data. It should be noted that in addition to processing the original subject image data to obtain the target subject visual effect image data, the trained large-parameter generation model can also be used as an independent visual effect model.
[0119] S203. Generate the training data of the visual effect model according to the target subject visual effect image data and the original subject image data.
[0120] The target subject visual effect image data is the gold standard data of the training data, and the original image data corresponding to the target subject visual effect image data is the input data of the model.
[0121] After obtaining the target subject visual effect image data, according to the target subject visual effect image data, determine the original subject image data corresponding to the target subject visual effect image data from the original subject image data, and then determine the target subject visual effect image data and the original subject image data corresponding to the target subject visual effect image data as the training data of the visual effect model.
[0122] Based on the generated training data, the small-parameter generation model can be trained to obtain the visual effect model. Since the target subject visual effect image data has a high consistency with the original subject image data, using the target subject visual effect image data for the small-parameter generation model will not interfere with the training of the small-parameter generation model, so that a lightweight visual effect processing model can be obtained and run on the intelligent terminal side, enabling the terminal user to experience rich visual effects, solving the pain point that a large amount of subject visual effect image data is required to implement visual effects on the terminal side, and greatly reducing the labor cost and time cost.
[0123] In the method for generating model training data provided by the embodiments of the present application, the original subject image data and the subject visual effect image data required by the visual effect model are used to train a preset large-parameter generation model to obtain a trained large-parameter generation model. Then, the original subject image data is processed by the trained large-parameter generation model to obtain the target subject visual effect image data. After that, based on the target subject visual effect image data and the original subject image data, the training data of the visual effect model is generated. Among them, the subject visual effect image data is obtained by performing visual effect processing on the original subject image data through a diffusion model. In this method, the original subject image data and the subject visual effect image data obtained by the diffusion model are used to train a preset large-parameter generation model. Due to the large number of parameters, the large-parameter generation model has a powerful learning ability, enabling it to learn more details and features of the original subject image data. Therefore, when performing visual effect processing, it can more accurately replicate the details of the original subject image data, and then the trained large-parameter generation model can generate a visual effect image with higher consistency with the original subject image data. Based on this, when the original subject image data is processed by the trained large-parameter generation model, the obtained target subject visual effect image data will have a high consistency with the original subject image data. In this way, using the target subject visual effect image data for model training is not likely to cause model training failure. And in order to train a lightweight visual effect model that can run on the terminal side, the target subject visual effect image data needs to be trained in a small-parameter model. Due to the high consistency between the target subject visual effect image data and the original subject image data, it will not interfere with the training of the small-parameter model, making it possible to deploy a lightweight visual effect model on the terminal side, and thus enabling terminal users to experience richer visual effects.
[0124] Based on the above embodiments, an embodiment is provided to illustrate the process of training a preset large-parameter generation model using the original subject image data and the subject visual effect image data required by the visual effect model to obtain a trained large-parameter generation model.
[0125] In an exemplary embodiment, as Figure 5 shown, training a preset large-parameter generation model using the original subject image data and the subject visual effect image data required by the visual effect model to obtain a trained large-parameter generation model includes:
[0126] S301, inputting the original subject image data into the large-parameter generation model to obtain a predicted subject visual effect image output by the large-parameter generation model.
[0127] When training a large-parameter generation model, the original body image data is input into the large-parameter generation model. The large-parameter generation model analyzes and processes the original body image data to obtain the visual effect map data corresponding to the original body image data predicted by the large-parameter generation model. Naturally, the visual effect map data corresponding to the predicted original body image data is the predicted body visual effect map.
[0128] S302. Determine the trained large-parameter generation model according to the predicted body visual effect map and the body visual effect map data.
[0129] After obtaining the predicted body visual effect map, based on the predicted body visual effect map and the body visual effect map data, adjust the large-parameter generation model to obtain the trained large-parameter generation model.
[0130] In the embodiments of the present application, the large-parameter generation model uses the body visual effect map data as the gold standard data for supervised learning. Based on this, after obtaining the predicted body visual effect map output by the large-parameter generation model, calculate the loss value of the large-parameter generation model according to the predicted body visual effect map and the body visual effect map data generated by the diffusion model, so as to adjust the large-parameter generation model based on the loss value to obtain the trained large-parameter generation model.
[0131] In the method for generating model training data provided by the embodiments of the present application, by inputting the original body image data into the large-parameter generation model, the predicted body visual effect map output by the large-parameter generation model is obtained. Then, according to the predicted body visual effect map and the body visual effect map data, the trained large-parameter generation model is determined. In this method, an optional way to train the large-parameter generation model is provided; by inputting the original body image data into the large-parameter generation model, the large-parameter generation model processes the original body image data to obtain the predicted body visual effect map. The predicted body visual effect map data is combined with the body visual effect map data generated by the diffusion model to continuously optimize the large-parameter generation model, enabling the large-parameter generation model to learn the ability of visual effect processing from the data generated by the diffusion model, laying a foundation for improving the consistency between the visual effect map and the original image.
[0132] Based on the above embodiments, an embodiment is provided to illustrate the process of determining the trained large-parameter generation model according to the predicted body visual effect map and the body visual effect map data.
[0133] In an exemplary embodiment, as Figure 6 shown, determining the trained large-parameter generation model according to the predicted body visual effect map and the body visual effect map data includes:
[0134] S401. Determine the prediction loss of the large-parameter generation model according to the predicted main body visual effect map, the main body visual effect map data, and a preset loss function.
[0135] In the embodiments of the present application, the preset loss function may include a perception loss, an adversarial loss, etc.
[0136] After obtaining the predicted main body visual effect map, the predicted main body visual effect map and the main body visual effect map data can be combined with the preset loss function to calculate the prediction loss of the large-parameter generation model.
[0137] For example, the predicted main body visual effect map and the main body visual effect map data can be substituted into the loss function to calculate the difference value between the predicted main body visual effect map and the main body visual effect map data. Naturally, the difference value between the predicted main body visual effect map and the main body visual effect map data is the prediction loss of the large-parameter generation model.
[0138] S402. Based on the prediction loss, adjust the model parameters of the large-parameter generation model until the training is completed to obtain the trained large-parameter generation model.
[0139] After obtaining the prediction loss of the large-parameter generation model, use the prediction loss to update the model parameters of the large-parameter generation model to reduce the loss between the predicted main body visual effect map and the main body visual effect map data, so that the predicted main body visual effect map generated by the large-parameter generation model approaches the main body visual effect map data, thereby achieving the learning of the visual effect processing ability and obtaining the trained large-parameter generation model.
[0140] After adjusting the model parameters of the large-parameter generation model according to the prediction loss, continue to train the large-parameter generation model with the current parameters, and then calculate the prediction loss of the large-parameter generation model until the prediction loss is less than the preset loss threshold, and stop training the large-parameter generation model to obtain the trained large-parameter generation model.
[0141] In the method for generating model training data provided by the embodiments of the present application, according to the predicted subject visual effect map, the subject visual effect map data, and a preset loss function, the prediction loss of the large-parameter generation model is determined. Then, based on the prediction loss, the model parameters of the large-parameter generation model are adjusted until the training is completed, and the trained large-parameter generation model is obtained. In this method, after obtaining the predicted subject visual effect map, the prediction loss of the large-parameter generation model is calculated by combining the subject visual effect map data and the preset loss function. The model parameters of the large-parameter generation model are adjusted through the prediction loss, so that the predicted value of the large-parameter generation model gradually approaches the real situation, and the accuracy of the prediction of the large-parameter generation model is continuously optimized, thereby improving the performance of the large-parameter generation model.
[0142] Based on any of the above embodiments, an embodiment is provided to illustrate the process of generating the training data of the visual effect model according to the target subject visual effect map data and the original subject image data.
[0143] In an exemplary embodiment, as Figure 7 shown, generating the training data of the visual effect model according to the target subject visual effect map data and the original subject image data includes:
[0144] S501, performing subject consistency screening and background consistency processing on the target subject visual effect map data to obtain the processed subject visual effect map data.
[0145] To further ensure the consistency of the subject in the subject visual effect map, after obtaining the target subject visual effect map data, subject consistency screening can be performed on the target subject visual effect map data to remove the data in the target subject visual effect map data that is significantly different from the original subject. Still taking face stylization as an example, the subject consistency screening here refers to face consistency screening, removing the data that is significantly different from the original face.
[0146] In addition, during the learning process of the visual effect processing of the large-parameter generation model, unnecessary visual effect processing is often also performed on the background information in the original image. Considering the limited representation and learning ability of the lightweight model, the background information in the target subject visual effect map data should be consistent with the background information in the original subject image data to exclude the interference caused by the visual effect processing of the background, so that the lightweight model can better focus on learning the visual effect processing of the foreground subject. Based on this, in the embodiments of the present application, after performing subject consistency screening on the target subject visual effect map data, background consistency processing is also performed.
[0147] In an embodiment of the present application, after obtaining the target subject visual effect map data, the target subject visual effect map data is subjected to subject consistency screening and background consistency, so as to remove the target subject visual effect map data in the target subject visual effect map data that has a large difference from the original subject, and process the background of the target subject visual effect map data to be consistent with the original subject data, so as to obtain the processed subject visual effect map data.
[0148] For example, a subject consistency screening algorithm can be used to perform subject consistency screening on the target subject visual effect map data, and a background fusion algorithm can be used to perform background consistency processing on the target subject visual effect map data to obtain the processed subject visual effect map data.
[0149] S502, Generate training data for the visual effect model according to the processed subject visual effect map data and the original subject data.
[0150] After obtaining the processed subject visual effect map data, based on the processed subject visual effect map data and the original subject data, generate training data for the visual effect model.
[0151] By performing subject consistency screening and background consistency processing on the target subject visual effect map data, the obtained processed subject visual effect map data has higher consistency with the original subject data. At this time, based on the processed subject visual effect map data and the original subject data, generating training data for the visual effect model, the visual effect maps in the training data have higher consistency with the original maps, which further improves the possibility of successful training of the lightweight visual effect model and the visual effect processing ability.
[0152] Exemplarily, according to the processed subject visual effect map data, the original subject data corresponding to the processed subject visual effect map data is screened out from the original subject data, and then the processed subject visual effect map data and the screened original subject data corresponding to the processed subject visual effect map data are used as training data.
[0153] In the method for generating model training data provided by the embodiments of the present application, after performing subject consistency screening and background consistency processing on the visual effect map data of the target subject, the processed visual effect map data of the subject is obtained. Then, based on the processed visual effect map data of the subject and the original subject image data, the training data of the visual effect model is generated. In this method, after obtaining the visual effect map data of the target subject, in order to further improve the consistency between the visual effect map data of the subject and the original subject image data, subject consistency screening and background consistency processing are also performed on the visual effect map data of the target subject. The obtained processed visual effect map data of the subject is more consistent with the original image in terms of the subject and is also consistent with the original image in terms of the background, avoiding interference with the training of lightweight models; moreover, it minimizes the manual intervention in the consistency processing process, reduces the new product launch threshold, and saves the new product launch cost.
[0154] Based on the above embodiments, an embodiment is provided to illustrate the process of performing subject consistency screening and background consistency processing on the visual effect map data of the target subject to obtain the processed visual effect map data of the subject.
[0155] In an exemplary embodiment, as Figure 8 shown, performing subject consistency screening and background consistency processing on the visual effect map data of the target subject to obtain the processed visual effect map data of the subject includes:
[0156] S601, according to the original subject image data and the visual effect map data of the target subject, removing the visual effect map data of the target subject in the visual effect map data of the target subject that has a large difference from the original subject in the original image, to obtain a subject consistency visual effect map set.
[0157] In the embodiments of the present application, when performing subject consistency screening on the visual effect map data of the target subject, the visual effect map data of the target subject in the visual effect map data of the target subject that has a large difference from the original subject in the original image is removed. Among them, it mainly involves the consistency on the subject, so it is considered to extract the subjects in both the visual effect map data of the target subject and the original subject image data.
[0158] Based on this, the subject of the original subject image data can be extracted to obtain the subject of the original subject image data, and the subject of the visual effect map data of the target subject can be extracted to obtain the subject of the visual effect map data of the target subject. Then, the subject of the original subject image data and the subject of the visual effect map data of the target subject are compared. If for a certain visual effect map data of the target subject, the subject of the visual effect map data of the target subject has a large difference from the subject of the corresponding original subject image data, then this visual effect map data of the target subject can be removed, and the remaining visual effect map data of the target subject constitutes the subject consistency visual effect map set.
[0159] Among them, when performing subject extraction on the original subject image data and the target subject visual effect image data, a subject extraction algorithm can be adopted. For example, an image segmentation algorithm based on graph cut can be used to separate the subject from the background.
[0160] S602. According to the original subject image data and the target subject visual effect image data, replace the visual effect background in the subject consistency visual effect image set with the original image background to obtain the processed target subject visual effect image data.
[0161] In the embodiment of the present application, for the background consistency processing of the target subject visual effect image data, the visual effect background in the subject consistency visual effect image set is replaced with the original image background. Among them, mainly the background is replaced, so the background of the original image can be considered to be fused with the foreground of the visual effect image, that is, the replacement of the background is realized.
[0162] Based on this, a background fusion algorithm can be adopted to process the original subject image data and the target subject visual effect image data, so as to fuse the background of the original subject image data with the foreground of the target subject visual effect image data to obtain the processed target subject visual effect image data.
[0163] In the method for generating model training data provided by the embodiment of the present application, according to the original subject image data and the target subject visual effect image data, the target subject visual effect images in the target subject visual effect image data that are significantly different from the original subject are removed to obtain a subject consistency visual effect image set. Then, according to the original subject image data and the target subject visual effect image data, the visual effect background in the subject consistency visual effect image set is replaced with the original image background to obtain the processed target subject visual effect image data. In this method, by removing the target subject visual effect images in the target subject visual effect image data that are significantly different from the original subject, the consistency of the visual effect image and the original image in terms of the subject is further improved. On this basis, the background of the visual effect image is replaced with the background of the original image, removing the visual effect processing effect of the background in the visual effect image, so that the processed target subject visual effect image data maintains high consistency with the original subject image data, providing data support for the rapid and accurate training of the lightweight visual effect model.
[0164] Based on the above embodiment, an embodiment is provided to illustrate the process of removing the target subject visual effect images in the target subject visual effect image data that are significantly different from the original subject according to the original subject image data and the target subject visual effect image data to obtain a subject consistency visual effect image set.
[0165] In an exemplary embodiment, as Figure 9 shown, according to the original subject image data and the target subject visual effect image data, removing the target subject visual effect images in the target subject visual effect image data that are significantly different from the original subject to obtain a subject consistency visual effect image set includes:
[0166] S701, perform subject segmentation on the original subject image data to obtain the original subject mask set; and, perform subject segmentation on the target subject visual effect image data to obtain the visual effect subject mask set.
[0167] Among them, the original subject mask set refers to the set of images obtained by performing subject segmentation on each original subject image in the original subject image data. The visual effect subject mask set refers to the set of images obtained by performing subject segmentation on each target subject visual effect image in the target subject visual effect image data.
[0168] When performing subject segmentation on the original subject image data, a subject segmentation algorithm can be used to process each original subject image in the original subject image data to obtain multiple processed original subject images, and these multiple processed original subject images are the original subject mask set.
[0169] Similarly, when performing subject segmentation on the target subject visual effect image data, a subject segmentation algorithm can also be used to process each target subject visual effect image in the target subject visual effect image data to obtain multiple processed target subject visual effect images, and these multiple processed target subject visual effect images are the visual effect subject mask set.
[0170] Taking the visual effect as face stylization as an example, the original subject image data is the original face image, the target subject visual effect image data is the target face stylization image. Use a face segmentation algorithm to perform face segmentation on the original face image to obtain the original face mask set, and use a face segmentation algorithm to perform face stylization on the target face stylization image to obtain the stylized face mask set.
[0171] S702, based on the original subject mask set and the visual effect subject mask set, perform image removal processing on the target subject visual effect image data to obtain the subject-consistent visual effect image set.
[0172] After obtaining the original subject mask set and the visual effect subject mask set, based on the original subject mask set and the visual effect subject mask set, the target subject visual effect images in the target subject visual effect image data that are significantly different from the original subject can be removed to obtain the subject-consistent visual effect image set.
[0173] In one embodiment, the original subject mask set and the visual effect subject mask set can be compared, and based on the comparison result, image removal processing is performed on the target subject visual effect image data. For example, for any target subject visual effect image data, if there is a large difference between the original subject mask corresponding to the target subject visual effect image data and the corresponding visual effect subject mask, then the target subject visual effect image data is removed.
[0174] In another embodiment, a difference evaluation index can be calculated based on the original subject mask set and the visual effect subject mask set, and the target subject visual effect map data can be processed for image removal according to the difference evaluation index. For example, for any target subject visual effect map data, the difference evaluation index of the target subject visual effect map data is calculated according to the original subject mask corresponding to the target subject visual effect map data and the corresponding visual effect subject mask. If the difference evaluation index of the target subject visual effect map data is greater than a preset threshold, the target subject visual effect map data is removed. Among them, the difference evaluation index can be the mean square error, or the difference evaluation index can also be selected as the peak signal-to-noise ratio, the structural similarity index, etc. When the difference evaluation index is the peak signal-to-noise ratio or the structural similarity index, at this time, it should be the target subject visual effect map data with the difference evaluation index less than the preset threshold that is removed.
[0175] In the method for generating model training data provided by the embodiments of the present application, the original subject map data is subjected to subject segmentation processing to obtain an original subject mask set; and, the target subject visual effect map data is subjected to subject segmentation processing to obtain a visual effect subject mask set, and then, according to the original subject mask set and the visual effect subject mask set, the target subject visual effect map data is processed for image removal to obtain a subject consistency visual effect map set. In this method, by performing subject segmentation processing on the original subject map data and the target subject visual effect map data respectively, the subject mask of the original subject map data and the subject mask of the target subject visual effect map data can be obtained, so that based on the difference between the two subject masks, the target subject visual effect maps in the target subject visual effect map data that are greatly different from the original subject can be quickly removed, improving the efficiency and accuracy of subject consistency screening.
[0176] Based on the above embodiments, an embodiment is provided to illustrate the process of processing the target subject visual effect map data for image removal according to the original subject mask set and the visual effect subject mask set to obtain a subject consistency visual effect map set.
[0177] In an exemplary embodiment, as Figure 10 shown, processing the target subject visual effect map data for image removal according to the original subject mask set and the visual effect subject mask set to obtain a subject consistency visual effect map set includes:
[0178] S801, determining the subject consistency decision data of each target subject visual effect map in the target subject visual effect map data according to the original subject mask set and the visual effect subject mask set.
[0179] Among them, the subject consistency decision data refers to the data used to determine whether there is a large difference between the target subject visual effect map data and the original subject map data.
[0180] Exemplarily, based on a preset calculation formula for the main body consistency decision data, the original main body mask set and the visual effect main body mask set can be substituted into the calculation to obtain the main body consistency decision data of each target main body visual effect map in the target main body visual effect map data.
[0181] Exemplarily, as Figure 11 shown, according to the original main body mask set and the visual effect main body mask set, determining the main body consistency decision data of each target main body visual effect map in the target main body visual effect map data includes:
[0182] S901. For any target main body visual effect map, based on the original main body mask set and the visual effect main body mask set, obtain the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask.
[0183] For any target main body visual effect map, first obtain the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask from the original main body mask set and the visual effect main body mask set, so as to calculate the main body consistency decision data of the target main body visual effect map based on the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask.
[0184] S902. Determine the overlapping area of the original main body mask and the visual effect main body mask.
[0185] After obtaining the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask, determine the overlapping area of the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask. For example, the overlapping area can be obtained by performing an intersection calculation on the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask.
[0186] S903. Determine the ratio of the overlapping area to the original main body mask as the main body consistency decision data of the target main body visual effect map.
[0187] After obtaining the overlapping area of the original main body mask corresponding to the target main body visual effect map and the corresponding visual effect main body mask, use the ratio of the overlapping area to the original main body mask as the main body consistency decision data of the target main body visual effect map.
[0188] Among them, the calculation formula of the main body consistency decision data is as shown in (1).
[0189]
[0190] Among them, M ori is the original main body mask; M style is the visual effect main body mask.
[0191] In this embodiment, for any target subject visual effect map, by obtaining the original subject mask and the corresponding visual effect subject mask corresponding to the target subject visual effect map from the original subject mask set and the visual effect subject mask set, the overlapping area between the original subject mask and the visual effect subject mask is determined, and finally the ratio of the overlapping area to the original subject mask is used as the subject consistency decision data for the target subject visual effect map. Equivalently, a calculation formula for determining the subject consistency decision data is provided, and based on this calculation formula, the subject consistency decision data for each target subject visual effect map can be quickly determined.
[0192] S802. According to the subject consistency decision data, perform image removal processing on the target subject visual effect map data to obtain a subject consistency visual effect map set.
[0193] After calculating the subject consistency decision data for each target subject visual effect map based on the above method, determine the removal object according to the subject consistency decision data of each target subject visual effect map to obtain a subject consistency visual effect map set.
[0194] Exemplarily, the subject consistency decision data of each target subject visual effect map can be compared with a preset threshold, and the target subject visual effect map with the subject consistency decision data less than the preset threshold is used as the removal object from the target subject visual effect map data, and the remaining target subject visual effect maps constitute the subject consistency visual effect map set.
[0195] Exemplarily, as Figure 12 shown, according to the subject consistency decision data, perform image removal processing on the target subject visual effect map data to obtain a subject consistency visual effect map set, including:
[0196] S1001. Obtain a preset subject consistency threshold.
[0197] Among them, the preset subject consistency threshold refers to a reference value for determining whether the subjects of the target subject visual effect map data and the original subject data are consistent. In the embodiment of the present application, the subject consistency threshold τ face can be 0.65. It should be noted that the subject consistency threshold can also be any other value determined according to actual needs, and the embodiment of the present application does not limit this.
[0198] In one embodiment, the method for obtaining the preset subject consistency threshold can be to obtain it based on a threshold setting interface. For example, when the subject consistency threshold needs to be obtained, a threshold setting interface can be displayed to the user, and the user can input the pre-calculated subject consistency threshold on the threshold setting interface. In this way, the server can obtain the subject consistency threshold.
[0199] In another embodiment, the way to obtain the preset subject consistency threshold may be to obtain the subject consistency threshold from historical subject consistency thresholds. For example, multiple historical subject consistency thresholds used at historical times are stored in a database. When it is necessary to obtain the subject consistency threshold, multiple historical subject consistency thresholds can be obtained from the database and presented to the user. The user can select one of the multiple historical subject consistency thresholds as the subject consistency threshold based on actual needs.
[0200] S1002, Remove the target subject visual effect map whose subject consistency decision data is less than the subject consistency threshold.
[0201] After obtaining the subject consistency threshold, compare the subject consistency decision data of each target subject visual effect map with the subject consistency threshold. For any target subject visual effect map, if the subject consistency decision data of the target subject visual effect map is less than the subject consistency threshold, then regard the target subject visual effect map as the object to be removed and remove it from the target subject visual effect map data.
[0202] S1003, Determine the remaining target subject visual effect maps in the target subject visual effect map data as the subject consistency visual effect map set.
[0203] After completing the image removal of the target subject visual effect map data based on the above-mentioned subject consistency threshold, the remaining target subject visual effect maps in the target subject visual effect map data are images with relatively small subject differences. Then, determine the remaining target subject visual effect maps in the target subject visual effect map data as the subject consistency visual effect map set.
[0204] In this embodiment, by obtaining the preset subject consistency threshold, it is possible to perform an image removal operation on the target subject visual effect map data based on the comparison result between the subject consistency threshold and the subject consistency decision data of each target subject visual effect map. The subject consistency threshold represents the reference value for whether the subject of the subject visual effect map and the subject original image is consistent. Then, the target subject visual effect map whose subject consistency decision data is less than the subject consistency threshold is the object to be removed, that is, remove the target subject visual effect map whose subject consistency decision data is less than the subject consistency threshold. The remaining target subject visual effect maps in the target subject visual effect map data are the subject consistency visual effect map set. By introducing the subject consistency threshold, it is possible to quickly determine the object to be removed from the target subject visual effect map data, providing an optional way to complete the rapid screening of subject consistency.
[0205] In the method for generating model training data provided by the embodiments of the present application, according to the original subject mask set and the visual effect subject mask set, the subject consistency decision data of each target subject visual effect map in the target subject visual effect map data is determined. According to each subject consistency decision data, image removal processing is performed on the target subject visual effect map data to obtain a subject consistency visual effect map set. In this method, the original subject mask set and the visual effect subject mask set are obtained through subject segmentation, and the subject consistency decision data of each target subject visual effect map in the target subject visual effect map data is calculated, so that based on this subject consistency visual effect map set, the data objects to be screened by the subject consistency screening can be quickly determined, providing an optional way for the rapid screening of subject consistency.
[0206] Based on the above embodiments, an embodiment is provided to illustrate the process of replacing the visual effect background in the subject consistency visual effect map set with the original map background according to the subject original map data and the target subject visual effect map data to obtain the processed subject visual effect map data.
[0207] In an exemplary embodiment, as Figure 13 shown, replacing the visual effect background in the subject consistency visual effect map set with the original map background according to the subject original map data and the target subject visual effect map data to obtain the processed subject visual effect map data includes:
[0208] S1101, performing subject segmentation processing on the subject original map data to obtain an original subject mask set; and performing subject segmentation processing on the target subject visual effect map data to obtain a visual effect subject mask set.
[0209] Consistent with the subject segmentation processing during the aforementioned subject consistency screening, a subject segmentation algorithm can be used to process each subject original map in the subject original map data to obtain multiple processed subject original maps, and these multiple processed subject original maps are the original subject mask set, and a subject segmentation algorithm is used to process each target subject visual effect map in the target subject visual effect map data to obtain multiple processed target subject visual effect maps, and these multiple processed target subject visual effect maps are the visual effect subject mask set.
[0210] S1102, determining the background consistency decision data of each subject consistency visual effect map in the subject consistency visual effect map set according to the original subject mask set and the visual effect subject mask set.
[0211] Among them, the background consistency decision data refers to the key data used for visual effect background replacement.
[0212] Exemplarily, based on a preset calculation formula for background consistency decision data, the original main body mask set and the visual effect main body mask set can be substituted into the calculation to obtain the background consistency decision data of each main body consistency visual effect map in the main body consistency visual effect map set.
[0213] Exemplarily, as Figure 14 shown, according to the original main body mask set and the visual effect main body mask set, determining the background consistency decision data of each main body consistency visual effect map in the main body consistency visual effect map set includes:
[0214] S1201, for any main body consistency visual effect map, according to the original main body mask set and the visual effect main body mask set, obtain the original main body mask corresponding to the main body consistency visual effect map and the corresponding visual effect main body mask.
[0215] For any main body consistency visual effect map, obtain the original main body mask corresponding to the main body consistency visual effect map and the corresponding visual effect main body mask from the original main body mask set and the visual effect main body mask set, so as to calculate the background consistency decision data of the main body consistency visual effect map based on the original main body mask and the corresponding visual effect main body mask corresponding to the main body consistency visual effect map.
[0216] S1202, determine the union mask of the original main body mask and the visual effect main body mask, and perform dilation operation and Gaussian blur processing on the union mask to obtain the processed union mask.
[0217] After obtaining the original main body mask and the corresponding visual effect main body mask corresponding to the main body consistency visual effect map, determine the union mask of the original main body mask and the corresponding visual effect main body mask corresponding to the main body consistency visual effect map. For example, perform a union calculation on the original main body mask and the corresponding visual effect main body mask corresponding to the main body consistency visual effect map to obtain the union mask. In this embodiment, calculating the union mask can ensure that the integrity of the main body part after visual effect processing will not be affected after fusing the original image background and the visual effect foreground.
[0218] Furthermore, after obtaining the union mask of the original main body mask and the visual effect main body mask, perform dilation operation and Gaussian blur processing on the union mask to make the edges smoother, which helps to more naturally realize the fusion of the original image background and the visual effect foreground.
[0219] S1203, determine the processed union mask as the background consistency decision data of the main body consistency visual effect map.
[0220] After obtaining the processed union mask, determine the processed union mask as the background consistency decision data of the main body consistency visual effect map.
[0221] In this embodiment, for any main body consistency visual effect map, the original main body mask and the corresponding visual effect main body mask corresponding to the main body consistency visual effect map are obtained from the original main body mask set and the visual effect main body mask set, and then the union mask of the original main body mask and the visual effect main body mask is determined. This union mask can ensure that the integrity of the main body part after visual effect processing will not be affected after the original image background and the visual effect foreground are fused. After performing dilation operation and Gaussian blur processing on the union mask, it is used as the background consistency decision data for the main body consistency visual effect map. Based on this background consistency decision data, background consistency processing is performed, so that the original image background and the visual effect foreground can be more naturally fused.
[0222] S1103. According to each background consistency decision data, perform visual effect background replacement on the main body consistency visual effect map set to obtain the processed main body visual effect map data.
[0223] After the background consistency decision data of each main body consistency visual effect map calculated based on the above method, based on the background consistency decision data of each main body consistency visual effect map, the visual effect foreground of the main body consistency visual effect map set is fused with the original image background of the main body original image data to obtain the processed main body visual effect map data.
[0224] Exemplarily, as Figure 15 shown, according to each background consistency decision data, perform visual effect background replacement on the main body consistency visual effect map set to obtain the processed main body visual effect map data, including:
[0225] S1301. Obtain the preset background replacement mathematical model and the main body consistency original map set corresponding to the main body consistency visual effect map set.
[0226] Among them, the preset background replacement mathematical model refers to the formula used for fusing the visual effect foreground and the original image background.
[0227] The background replacement mathematical model can be pre-stored in the mathematical model database. When the background replacement mathematical model needs to be obtained, based on the identification information of the background replacement mathematical model, the data model corresponding to the identification information can be matched from the mathematical model database, and the matched mathematical model is determined as the background replacement mathematical model.
[0228] The main body consistency original map set corresponding to the main body consistency visual effect map set can be obtained according to the main body original image data. Exemplarily, the main body original images corresponding to each main body consistency visual effect map in the main body consistency visual effect map set are obtained from the main body original image data, and the main body original images corresponding to each main body consistency visual effect map are used as the main body consistency original map set corresponding to the main body consistency visual effect map set.
[0229] S1302. According to each background consistency decision data, the background replacement mathematical model, and the main body consistency original atlas, perform visual effect background replacement on the main body consistency visual effect atlas to obtain the processed main body visual effect map data.
[0230] After obtaining the background replacement mathematical model and the main body consistency original atlas corresponding to the main body consistency visual effect atlas, each background consistency decision data and the main body consistency original atlas can be input into the background replacement mathematical model to perform visual effect background replacement on the main body consistency visual effect atlas, thereby obtaining the processed main body visual effect map data.
[0231] In the embodiment of the present application, the background replacement mathematical model can be as shown in the following formula (2).
[0232] I merge =(1 - M front )·I ori +M front ·I GAN (2)
[0233] Wherein, I merge is the processed main body visual effect map data; M front is the background consistency decision data; I ori is the main body consistency original atlas; I GAN is the main body consistency visual effect atlas.
[0234] For any main body consistency visual effect map, input the background consistency decision data corresponding to the main body consistency visual effect map and the corresponding main body consistency original map into the above background replacement mathematical model, and then the visual effect foreground of the main body consistency visual effect map can be fused with the original map background of the main body consistency original map to obtain the processed main body consistency visual effect map. After obtaining the processed main body consistency visual effect map for each main body consistency visual effect map in this way, all the processed main body consistency visual effect maps are the processed main body visual effect map data.
[0235] In this embodiment, by obtaining the preset background replacement mathematical model and the main body consistency original atlas corresponding to the main body consistency visual effect atlas, based on this background replacement mathematical model, the main body consistency original atlas and the main body consistency visual effect atlas corresponding to the main body consistency visual effect atlas can be input into the background replacement mathematical model to quickly obtain the processed main body visual effect map data; by introducing the background replacement mathematical model, quickly fuse the visual effect foreground of the main body consistency visual effect map with the original map background of the main body consistency original map, providing an optional way for realizing background consistency processing.
[0236] In the method for generating model training data provided by the embodiments of the present application, the original subject image data is subjected to subject segmentation processing to obtain an original subject mask set; and the target subject visual effect image data is subjected to subject segmentation processing to obtain a visual effect subject mask set; according to the original subject mask set and the visual effect subject mask set, background consistency decision data for each subject consistency visual effect image in the subject consistency visual effect image set is determined; according to each background consistency decision data, visual effect background replacement is performed on the subject consistency visual effect image set to obtain processed subject visual effect image data. In this method, by performing subject segmentation processing on the original subject image data and the target subject visual effect image data respectively, the subject mask of the original subject image data and the subject mask of the target subject visual effect image data can be obtained, so that the fusion of the visual effect foreground and the original image background can be quickly realized based on the set of the two subject masks, and the efficiency and accuracy of background consistency processing are improved.
[0237] Based on the above embodiments, an embodiment for the process of obtaining the original subject image data and the subject visual effect image data is provided for illustration.
[0238] In an exemplary embodiment, as Figure 16 shown, the method further includes:
[0239] S1401, obtain a plurality of historical original subject images.
[0240] In one embodiment, the way to obtain a plurality of historical original subject images is to obtain a plurality of historical original subject images from a database. A large number of historical original subject images are stored in the database. Some of these historical original subject images can be selected as a plurality of historical original subject images, or all the historical original subject images in the database can be used as a plurality of historical original subject images.
[0241] In another embodiment, the way to obtain a plurality of historical original subject images is to obtain a plurality of historical original subject images through a search engine or a professional historical image website. For example, a plurality of historical original subject images can be obtained in search engines such as Baidu, Google, and Sogou.
[0242] S1402, input each historical original subject image into a diffusion model to obtain a plurality of subject visual effect images output by the diffusion model.
[0243] After obtaining a plurality of historical original subject images, input each historical original subject image into a diffusion model, and the diffusion model performs visual effect processing on each historical original subject image to obtain a plurality of subject visual effect images output by the diffusion model.
[0244] S1403, determine each historical original subject image as the original subject image data, and determine each subject visual effect image as the subject visual effect image data.
[0245] Each obtained historical original subject image can be used as the original subject image data, and multiple subject visual effect images obtained through processing by the diffusion model can be used as the subject visual effect image data.
[0246] After obtaining each historical original subject image, preprocessing can also be performed on each historical original subject image. For example, images with low clarity in each historical original subject image can be deleted to prevent interference with the accuracy of subsequent visual effect model training. Alternatively, image enhancement technology can also be used to process each historical original subject image to improve the image quality of each historical original subject image.
[0247] In the method for generating model training data provided by the embodiments of the present application, by obtaining multiple historical original subject images, and then inputting each historical original subject image into the diffusion model, multiple subject visual effect images output by the diffusion model are obtained. After that, each historical original subject image is determined as the original subject image data, and each subject visual effect image is determined as the subject visual effect image data. In this method, an optional way to quickly obtain the original subject image data and the subject visual effect image data is provided; by obtaining multiple historical original subject images, data support is provided for the original subject image data. Inputting each obtained historical original subject image into the diffusion model, and performing visual effect processing through the diffusion model to obtain the subject visual effect image data, which lays a foundation for the training of the large-parameter generation model.
[0248] Based on the above embodiments, an embodiment is provided to illustrate the process before training the preset large-parameter generation model using the original subject image data and the subject visual effect image data to obtain the trained large-parameter generation model.
[0249] In an exemplary embodiment, as Figure 17 shown, before training the preset large-parameter generation model using the original subject image data and the subject visual effect image data to obtain the trained large-parameter generation model, the method further includes:
[0250] S1501, obtaining the original subject mask set of the original subject image data and the visual effect subject mask set of the subject visual effect image data.
[0251] As described above, the subject segmentation algorithm can be used to process each original subject image in the original subject image data to obtain multiple processed original subject images, and these multiple processed original subject images are the original subject mask set of the original subject image data. Using the subject segmentation algorithm to process each subject visual effect image in the subject visual effect image data to obtain multiple processed subject visual effect images, and these multiple processed subject visual effect images are the visual effect subject mask set of the subject visual effect image data.
[0252] S1502. Based on the original main body mask set and the visual effect main body mask set, perform main body consistency screening on the original main body image data and the main body visual effect image data to obtain the screened original main body image data and the screened main body visual effect image data.
[0253] In the embodiments of the present application, when performing main body consistency screening on the original main body image data and the main body visual effect image data based on the original main body mask set and the visual effect main body mask set, the main body consistency decision data of each main body visual effect image in the main body visual effect image data can also be calculated using formula (1). Then, based on the main body consistency decision data of each main body visual effect image, remove the main body visual effect images whose main body consistency decision data is less than the main body consistency threshold. The remaining main body visual effect images are the screened main body visual effect image data. Obtain the original main body image data corresponding to the screened main body visual effect image data from the original main body image data and use it as the screened original main body image data.
[0254] In the method for generating model training data provided by the embodiments of the present application, by obtaining the original main body mask set of the original main body image data and the visual effect main body mask set of the main body visual effect image data, and then performing main body consistency screening on the original main body image data and the main body visual effect image data based on the original main body mask set and the visual effect main body mask set, the screened original main body image data and the screened main body visual effect image data are obtained. In this method, before the large-parameter generation model, perform a main body consistency screening on the original main body mask set and the visual effect main body mask set using the main body mask corresponding to the original main body image data and the main body mask corresponding to the main body visual effect image data, so that the consistency of the input data of the large-parameter generation model is higher, improving the performance of the large-parameter generation model, and further improving the reliability of the training data.
[0255] Based on the above embodiments, an embodiment is provided to illustrate the training process of the above visual effect model.
[0256] In an exemplary embodiment, as Figure 18 shown, the method further includes:
[0257] S1601. Based on the target main body visual effect image data and the original main body image data, obtain the target original main body image data.
[0258] The target original main body image data refers to the original image corresponding to the target main body visual effect image data.
[0259] Exemplarily, based on the target main body visual effect image data, obtain the original main body image corresponding to each target main body visual effect image from the original main body image data, and then determine the original main body image corresponding to each target main body visual effect image as the target original main body image data.
[0260] S1602. Input the original image data of the target object into a pre-set small-parameter generation model to obtain the predicted visual effect map data of the object output by the small-parameter generation model.
[0261] As described above, the visual effect model is a lightweight model that runs on the intelligent terminal side for visual effect processing. Therefore, when training the visual effect model, the initial model selected should be a small-parameter generation model to achieve the purpose of lightweight deployment. In the embodiments of the present application, the small-parameter generation model can be a generative adversarial network with a small number of parameters, and the number of parameters can be about 1M.
[0262] After obtaining the original image data of the target object, input the original image data of the target object into a pre-set small-parameter generation model. The small-parameter generation model processes the original image data of the target object to obtain the visual effect map data corresponding to the original image data of the target object predicted by the small-parameter generation model. Naturally, the visual effect map data corresponding to the predicted original image data of the target object is the predicted visual effect map data of the object.
[0263] S1603. According to the predicted visual effect map data of the object, the visual effect map data of the target object, and a pre-set loss function, adjust the model parameters of the small-parameter generation model until the training is completed to obtain the visual effect model.
[0264] After obtaining the predicted visual effect map data of the object, based on the predicted visual effect map data of the object, the visual effect map data of the target object, and a pre-set loss function, adjust the small-parameter generation model to obtain the visual effect model. In the embodiments of the present application, the pre-set loss function can include perceptual loss and adversarial loss, etc.
[0265] Taking the visual effect map data of the target object as the gold standard data of the small-parameter generation model, input the predicted visual effect map data of the object and the visual effect map data of the target object into the loss function to calculate the difference value between the predicted visual effect map data of the object and the visual effect map data of the target object. Naturally, the difference value between the predicted visual effect map data of the object and the visual effect map data of the target object is the prediction loss of the small-parameter generation model.
[0266] Furthermore, through the prediction loss, update the model parameters of the small-parameter generation model to reduce the loss between the predicted visual effect map data of the object and the visual effect map data of the target object. When the prediction loss is less than the pre-set loss threshold, stop training the small-parameter generation model to obtain the visual effect model.
[0267] In the method for generating model training data provided by the embodiments of the present application, according to the target subject visual effect map data and the original subject image data, the original target subject image data is obtained. Then, the original target subject image data is input into a preset small-parameter generation model to obtain the predicted subject visual effect map data output by the small-parameter generation model. After that, according to the predicted subject visual effect map data, the target subject visual effect map data, and a preset loss function, the model parameters of the small-parameter generation model are adjusted until the training is completed, and a visual effect model is obtained. In this method, the small-parameter generation model is used as the initial model of the visual effect model. The original target subject image data is input into the small-parameter generation model to obtain a predicted value, and the target subject visual effect map data with high consistency is used as the gold standard data. Based on the loss between the predicted value and the gold standard, the model parameters of the small-parameter generation model are adjusted to continuously optimize the prediction accuracy of the small-parameter generation model, so as to obtain a visual effect model with better performance, laying a foundation for the lightweight model deployment on the terminal side.
[0268] Based on the above embodiments, an embodiment of the usage process of the above visual effect model is provided for illustration.
[0269] In an exemplary embodiment, as Figure 19 shown, the method further includes:
[0270] S1701, obtain the original subject image to be processed.
[0271] The original subject image to be processed is the original subject image that the user currently needs to process.
[0272] S1702, input the original subject image into the visual effect model to obtain the subject visual effect image of the original subject image.
[0273] Based on the above training, the visual effect model can be obtained. The original subject image to be processed can be input into the visual effect model, and the visual effect model can perform visual effect processing on the original subject image to be processed to obtain the subject visual effect image of the original subject image.
[0274] After the visual effect model is trained, it can be deployed on the terminal side for use. At this time, the original subject image to be processed can be the original subject image currently captured by the terminal user, or the original subject image uploaded from the terminal image library.
[0275] Taking the visual effect of face stylization as an example, the terminal user can perform face stylization on the original face image to be processed through the deployed visual effect model to obtain the face stylized image of the original face image.
[0276] In the method for generating model training data provided by the embodiments of the present application, by obtaining the original subject image to be processed, and then inputting the original subject image into the visual effect model, the subject visual effect image of the original subject image is obtained. In this method, by obtaining the original subject image to be processed and performing visual effect processing on the original subject image through the trained visual effect model, the subject visual effect image of the original subject image can be quickly obtained. The visual effect model is trained based on the subject visual effect map data with high consistency, so that the visual effect model has strong visual effect processing ability and improves the consistency between the visual effect map and the original image.
[0277] In addition, in an exemplary embodiment, the present application also provides a method for generating model training data, as Figure 20 shown, which may include the following steps:
[0278] S1801, input the original subject image data into the large-parameter generation model to obtain the predicted subject visual effect map output by the large-parameter generation model.
[0279] S1802, determine the prediction loss of the large-parameter generation model according to the predicted subject visual effect map, the subject visual effect map data, and the preset loss function.
[0280] Among them, the subject visual effect map data is obtained by performing visual effect processing on the original subject image data through the diffusion model.
[0281] S1803, based on the prediction loss, adjust the model parameters of the large-parameter generation model until the training is completed to obtain the trained large-parameter generation model.
[0282] S1804, process the original subject image data according to the trained large-parameter generation model to obtain the target subject visual effect map data.
[0283] S1805, perform subject consistency screening and background consistency processing on the target subject visual effect map data to obtain the processed subject visual effect map data.
[0284] Exemplarily, according to the original subject image data and the target subject visual effect map data, remove the target subject visual effect maps in the target subject visual effect map data that have a large difference from the original image subject to obtain a subject consistency visual effect map set; according to the original subject image data and the target subject visual effect map data, replace the visual effect background in the subject consistency visual effect map set with the original image background to obtain the processed subject visual effect map data.
[0285] S1806, obtain the target original subject image data according to the target subject visual effect map data and the original subject image data.
[0286] S1807. Input the original image data of the target subject into a pre-set small-parameter generation model to obtain the predicted subject visual effect image data output by the small-parameter generation model.
[0287] S1808. According to the predicted subject visual effect image data, the target subject visual effect image data, and a pre-set loss function, adjust the model parameters of the small-parameter generation model until the training is completed to obtain a visual effect model.
[0288] S1809. Obtain the original subject image to be processed.
[0289] S1810. Input the original subject image into the visual effect model to obtain the subject visual effect image of the original subject image.
[0290] The processes of S1801 - S1810 above can refer to the description of the above method embodiments, and their implementation principles and technical effects are similar, so they will not be elaborated here.
[0291] In addition, in an exemplary embodiment, as Figure 21 shown, taking subject stylization as an example, the construction process of a lightweight subject stylization model will be described.
[0292] (1) Perform subject consistency screening on the DM stylized dataset generated by the DM model to obtain the screened DM stylized dataset. Among them, the subject in the image is a puppy. The screened DM stylized dataset includes multiple DM stylized images, providing data for the training of the large-parameter GAN.
[0293] (2) Input the original images corresponding to the screened DM stylized dataset into the large-parameter GAN model to generate predicted GAN-generated images, and use the DM stylized images in the screened DM stylized dataset as the gold standard for supervised learning. After the training is completed, the large-parameter GAN model already has the subject stylization ability equivalent to that of the DM model. At this time, use the trained large-parameter GAN model to perform subject stylization on the original images to generate corresponding stylized data.
[0294] (3) Perform subject consistency screening and background consistency processing on the stylized data generated by the trained large-parameter GAN model to obtain highly consistent stylized images.
[0295] (4) Input the original images corresponding to the highly consistent stylized images into the small-parameter GAN model, and use the highly consistent stylized images as the gold standard for supervised learning to obtain the lightweight GAN generation result. The trained small-parameter GAN model can be used as a lightweight subject stylization model and deployed on the terminal side for users to use.
[0296] In this embodiment, the main body visual effect map data obtained from the main body original map data and the diffusion model is used to train a preset large-parameter generation model. Due to the large number of parameters, the large-parameter generation model has a powerful learning ability, enabling it to learn more details and features of the main body original map data. Therefore, when performing visual effect processing, it can more accurately replicate the details of the main body original map data, and further enable the trained large-parameter generation model to generate a visual effect map with higher consistency with the main body original map data. Based on this, by processing the main body original map data with the trained large-parameter generation model, the obtained target main body visual effect map data will have a high consistency with the main body original map data. In this way, using the target main body visual effect map data for model training is not likely to cause model training failure; and in order to be able to train a lightweight visual effect model that runs on the terminal side, it is necessary to train the target main body visual effect map data in a small-parameter model. Due to the high consistency between the target main body visual effect map data and the main body original map data, it will not interfere with the training of the small-parameter model, making it possible to deploy a lightweight visual effect model on the terminal side, and further enabling terminal users to experience richer visual effects.
[0297] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0298] Based on the same inventive concept, an embodiment of the present application also provides a model training data generation device for implementing the above-mentioned model training data generation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following model training data generation device embodiments can refer to the limitations on the model training data generation method in the above text, and will not be repeated here.
[0299] In an exemplary embodiment, as Figure 22 shown, a model training data generation device 1 is provided, including: a model training module 10, a data processing module 20, and a training data generation module 30, where:
[0300] The model training module 10 is used to train a preset large-parameter generation model with the original subject image data and the subject visual effect image data required by the visual effect model, and obtain a trained large-parameter generation model; the subject visual effect image data is obtained by performing visual effect processing on the original subject image data through a diffusion model;
[0301] The data processing module 20 is used to process the original subject image data according to the trained large-parameter generation model to obtain target subject visual effect image data;
[0302] The training data generation module 30 is used to generate training data for the visual effect model according to the target subject visual effect image data and the original subject image data.
[0303] In one embodiment, the above-mentioned model training module 10 is further used for:
[0304] Input the original subject image data into the large-parameter generation model to obtain a predicted subject visual effect image output by the large-parameter generation model; determine the trained large-parameter generation model according to the predicted subject visual effect image and the subject visual effect image data.
[0305] In one embodiment, the above-mentioned model training module 10 is further used for:
[0306] Determine the prediction loss of the large-parameter generation model according to the predicted subject visual effect image, the subject visual effect image data and a preset loss function; based on the prediction loss, adjust the model parameters of the large-parameter generation model until the training is completed to obtain a trained large-parameter generation model.
[0307] In one embodiment, the above-mentioned training data generation module 30 is further used for:
[0308] Perform subject consistency screening and background consistency processing on the target subject visual effect image data to obtain processed subject visual effect image data; generate training data for the visual effect model according to the processed subject visual effect image data and the original subject image data.
[0309] In one embodiment, the above-mentioned training data generation module 30 is further used for:
[0310] According to the original subject image data and the target subject visual effect image data, remove the target subject visual effect images in the target subject visual effect image data that have a large difference from the original subject to obtain a subject consistency visual effect image set; replace the visual effect background in the subject consistency visual effect image set with the original background to obtain processed subject visual effect image data.
[0311] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0312] According to the original subject image data and the target subject visual effect image data, remove the target subject visual effect images in the target subject visual effect image data that have a large difference from the original subject in the original image, and obtain a set of subject-consistent visual effect images, including:
[0313] Perform subject segmentation processing on the original subject image data to obtain a set of original subject masks; and perform subject segmentation processing on the target subject visual effect image data to obtain a set of visual effect subject masks; according to the set of original subject masks and the set of visual effect subject masks, perform image removal processing on the target subject visual effect image data to obtain a set of subject-consistent visual effect images.
[0314] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0315] According to the set of original subject masks and the set of visual effect subject masks, determine the subject consistency decision data of each target subject visual effect image in the target subject visual effect image data; according to the subject consistency decision data, perform image removal processing on the target subject visual effect image data to obtain a set of subject-consistent visual effect images.
[0316] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0317] For any target subject visual effect image, according to the set of original subject masks and the set of visual effect subject masks, obtain the corresponding original subject mask and the corresponding visual effect subject mask of the target subject visual effect image; determine the overlapping area of the original subject mask and the visual effect subject mask; determine the ratio of the overlapping area to the original subject mask as the subject consistency decision data of the target subject visual effect image.
[0318] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0319] Obtain a preset subject consistency threshold; remove the target subject visual effect images whose subject consistency decision data is less than the subject consistency threshold; determine the remaining target subject visual effect images in the target subject visual effect image data as a set of subject-consistent visual effect images.
[0320] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0321] Perform subject segmentation on the original subject image data to obtain the original subject mask set; and, perform subject segmentation on the target subject visual effect image data to obtain the visual effect subject mask set; determine the background consistency decision data for each subject consistency visual effect image in the subject consistency visual effect image set according to the original subject mask set and the visual effect subject mask set; perform visual effect background replacement on the subject consistency visual effect image set according to the respective background consistency decision data to obtain the processed subject visual effect image data.
[0322] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0323] For any subject consistency visual effect image, obtain the corresponding original subject mask and the corresponding visual effect subject mask of the subject consistency visual effect image according to the original subject mask set and the visual effect subject mask set; determine the union mask of the original subject mask and the visual effect subject mask, and perform dilation operation and Gaussian blur processing on the union mask to obtain the processed union mask; determine the processed union mask as the background consistency decision data of the subject consistency visual effect image.
[0324] In one embodiment, the above-mentioned training data generation module 30 is further configured to:
[0325] Obtain a preset background replacement mathematical model and the subject consistency original image set corresponding to the subject consistency visual effect image set; perform visual effect background replacement on the subject consistency visual effect image set according to the respective background consistency decision data, the background replacement mathematical model, and the subject consistency original image set to obtain the processed subject visual effect image data.
[0326] In one embodiment, the above-mentioned model training data generation device 1 further includes:
[0327] A historical original image acquisition module, configured to acquire a plurality of historical subject original images;
[0328] A visual effect processing module, configured to input each historical subject original image into a diffusion model to obtain a plurality of subject visual effect images output by the diffusion model;
[0329] A data determination module, configured to determine each historical subject original image as subject original image data, and determine each subject visual effect image as subject visual effect image data.
[0330] In one embodiment, the above-mentioned model training data generation device 1 further includes:
[0331] A mask acquisition module, which acquires the original subject mask set of the subject original image data and the visual effect subject mask set of the subject visual effect image data;
[0332] An image screening module, configured to perform subject consistency screening on the original subject image data and the subject visual effect image data according to the original subject mask set and the visual effect subject mask set, so as to obtain the screened original subject image data and the screened subject visual effect image data.
[0333] In one embodiment, the above-mentioned generating device 1 of model training data further includes:
[0334] An original image data acquisition module, configured to acquire target original subject image data according to the target subject visual effect image data and the original subject image data;
[0335] Input the target original subject image data into a preset small-parameter generation model to obtain the predicted subject visual effect image data output by the small-parameter generation model;
[0336] According to the predicted subject visual effect image data, the target subject visual effect image data and a preset loss function, adjust the model parameters of the small-parameter generation model until the training is completed to obtain a visual effect model.
[0337] In one embodiment, the above-mentioned generating device 1 of model training data further includes:
[0338] An image to be processed acquisition module, configured to acquire an original subject image to be processed;
[0339] A visual effect image determination module, configured to input the original subject image into the visual effect model to obtain the subject visual effect image of the original subject image.
[0340] Each module in the above-mentioned generating device of model training data can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0341] It should be noted that the personnel information (including but not limited to personnel device information, personnel personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the personnel or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations.
[0342] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0343] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0344] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for generating model training data, characterized in that: The method comprises: The main original image data and the main visual special effect image data required by the visual special effect model are used to train the preset large parameter generation model to obtain the trained large parameter generation model; the main visual special effect image data is obtained by performing visual special effect processing on the main original image data through the diffusion model; Processing the subject original image data according to the trained large parameter generation model to obtain target subject visual special effect image data; The training data of the visual effects model is generated according to the target subject visual effects image data and the subject original image data.
2. The method according to claim 1, characterized in that The method uses the subject original image data and the subject visual effect image data required by the visual effect model to train the preset large parameter generation model to obtain the trained large parameter generation model, including: Inputting the subject original image data into the large parameter generation model to obtain a predicted subject visual special effects image output by the large parameter generation model; The trained large-parameter generation model is determined based on the predicted subject visual effects map and the subject visual effects map data.
3. The method according to claim 2, characterized in that The step of determining the trained large-parameter generation model based on the predicted subject visual special effects graph and the subject visual special effects graph data comprises: Determining the prediction loss of the large-parameter generation model according to the predicted subject visual effects graph, the subject visual effects graph data and a preset loss function; Based on the prediction loss, the model parameters of the large parameter generation model are adjusted until the training is completed to obtain the trained large parameter generation model.
4. The method according to any one of claims 1 to 3, characterized in that: The step of generating training data for the visual effects model according to the target subject visual effects image data and the subject original image data comprises: Performing subject consistency screening and background consistency processing on the target subject visual special effects image data to obtain processed subject visual special effects image data; The training data of the visual effects model is generated according to the processed subject visual effects image data and the subject original image data.
5. The method according to claim 4, characterized in that The subject consistency screening and background consistency processing of the target subject visual special effects image data to obtain the processed subject visual special effects image data includes: According to the subject original image data and the target subject visual special effect image data, the target subject visual special effect images in the target subject visual special effect image data that are greatly different from the original subject image are removed to obtain a subject consistent visual special effect image set; According to the subject original image data and the target subject visual special effect image data, the visual special effect background in the subject consistent visual special effect image set is replaced with the original image background to obtain the processed subject visual special effect image data.
6. The method according to claim 5, characterized in that The method of removing the target subject visual effects image data that is greatly different from the original subject image in the target subject visual effects image data according to the subject original image data and the target subject visual effects image data, to obtain a subject consistent visual effects image set, comprises: Performing subject segmentation processing on the subject original image data to obtain an original subject mask set; and performing subject segmentation processing on the target subject visual effect image data to obtain a visual effect subject mask set; According to the original subject mask set and the visual effect subject mask set, image removal processing is performed on the target subject visual effect map data to obtain the subject consistency visual effect map set.
7. The method according to claim 6, characterized in that The step of performing image removal processing on the target subject visual effect graph data according to the original subject mask set and the visual effect subject mask set to obtain the subject consistency visual effect graph set includes: Determining subject consistency decision data of each target subject visual special effect graph in the target subject visual special effect graph data according to the original subject mask set and the visual special effect subject mask set; According to each of the subject consistency decision data, image removal processing is performed on the target subject visual special effects map data to obtain the subject consistency visual special effects map set.
8. The method according to claim 7, characterized in that The determining, according to the original subject mask set and the visual effect subject mask set, subject consistency decision data of each target subject visual effect graph in the target subject visual effect graph data comprises: For any target subject visual effect image, according to the original subject mask set and the visual effect subject mask set, obtain the original subject mask and the corresponding visual effect subject mask corresponding to the target subject visual effect image; Determining an overlapping area between the original main body mask and the visual effect main body mask; The ratio of the overlapped area to the original subject mask is determined as subject consistency decision data of the target subject visual special effect image.
9. The method according to claim 7, characterized in that: The step of performing image removal processing on the target subject visual special effect graph data according to each subject consistency decision data to obtain the subject consistency visual special effect graph set includes: Get the preset subject consistency threshold; Remove the target subject visual special effect image whose subject consistency decision data is less than the subject consistency threshold; The remaining target subject visual effects graphs in the target subject visual effects graph data are determined as the subject consistency visual effects graph set.
10. The method according to claim 5, characterized in that The step of replacing the visual effects background in the subject consistency visual effects picture set with the original picture background according to the subject original picture data and the target subject visual effects picture data to obtain the processed subject visual effects picture data includes: Performing subject segmentation processing on the subject original image data to obtain an original subject mask set; and performing subject segmentation processing on the target subject visual effect image data to obtain a visual effect subject mask set; Determining background consistency decision data for each subject consistency visual effect graph in the subject consistency visual effect graph set according to the original subject mask set and the visual effect subject mask set; According to each of the background consistency decision data, the visual effects background is replaced on the subject consistency visual effects atlas to obtain the processed subject visual effects image data.
11. The method according to claim 10, characterized in that The determining, according to the original subject mask set and the visual effect subject mask set, background consistency decision data of each subject consistency visual effect graph in the subject consistency visual effect graph set comprises: For any subject-consistent visual effect graph, obtaining an original subject mask and a corresponding visual effect subject mask corresponding to the subject-consistent visual effect graph according to the original subject mask set and the visual effect subject mask set; Determine a union mask of the original main body mask and the visual effect main body mask, and perform a dilation operation and a Gaussian blur process on the union mask to obtain a processed union mask; The processed union mask is determined as background consistency decision data of the subject consistency visual effect map.
12. The method according to claim 10, characterized in that The step of performing visual effects background replacement on the subject consistency visual effects atlas according to each of the background consistency decision data to obtain the processed subject visual effects atlas data includes: Obtaining a preset background replacement mathematical model and a subject consistency original atlas corresponding to the subject consistency visual special effects atlas; According to each of the background consistency decision data, the background replacement mathematical model and the subject consistency original atlas, the subject consistency visual special effects atlas is subjected to visual special effects background replacement to obtain the processed subject visual special effects atlas data.
13. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Obtain original images of multiple historical subjects; Inputting each of the historical subject original images into the diffusion model to obtain a plurality of subject visual special effect images output by the diffusion model; Each of the historical main original images is determined as the main original image data, and each of the main visual special effect images is determined as the main visual special effect image data.
14. The method according to any one of claims 1 to 3, characterized in that: Before using the subject original image data and the subject visual special effect image data to train a preset large parameter generation model to obtain a trained large parameter generation model, the method further includes: Acquire an original subject mask set of the subject original image data and a visual effect subject mask set of the subject visual effect image data; According to the original subject mask set and the visual special effect subject mask set, the subject original image data and the subject visual special effect image data are screened for subject consistency to obtain screened subject original image data and screened subject visual special effect image data.
15. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Acquire the target subject original image data according to the target subject visual special effect image data and the subject original image data; Inputting the target subject original image data into a preset small parameter amount generation model to obtain the predicted subject visual special effect image data output by the small parameter amount generation model; According to the predicted subject visual special effects image data, the target subject visual special effects image data and a preset loss function, the model parameters of the small parameter generation model are adjusted until the training is completed to obtain the visual special effects model.
16. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Obtaining an original subject image to be processed; The original subject image is input into the visual effects model to obtain a subject visual effects image of the original subject image.
17. A device for generating model training data, characterized in that: The device comprises: A model training module is used to train a preset large-parameter generation model using the subject original image data and subject visual effect image data required by the visual effect model to obtain a trained large-parameter generation model; the subject visual effect image data is obtained by performing visual effect processing on the subject original image data through a diffusion model; A data processing module, used for processing the subject original image data according to the trained large parameter generation model to obtain target subject visual special effect image data; The training data generation module is used to generate the training data of the visual effects model according to the target subject visual effects image data and the subject original image data.
18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 16 are implemented.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.
20. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.