Training of multimedia resource generation model and multimedia resource generation method
By generating and utilizing positive and negative sample resource pairs for iterative training, the problem of poor quality of diffusion model training samples is solved, and the denoising performance and generation effect of the multimedia resource generation model are improved.
Patent Information
- Application Number
- CN202510983256.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-17
AI Technical Summary
During the existing diffusion model training process, the quality of training samples is poor, resulting in poor model training results, and then poor denoising performance and generation effect of the multimedia resource generation model.
By obtaining resource generation description information and preset empty description information, and using preset guidance parameters and scaled guidance parameters, positive and negative sample multimedia resources are generated, and positive and negative sample resource pairs are constructed. The multimedia resource generation model is iteratively trained, and the denoising performance is controlled in combination with the temperature coefficient.
It improves the model training efficiency and multimedia resource generation effect, solves various quality defects of diffusion model generation resources in a targeted manner, and improves denoising performance.
Smart Images

Figure CN120494016B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a multimedia resource generation model training and multimedia resource generation method. Background Art
[0002] With the rapid development of artificial intelligence (AI), generative models such as diffusion models have gained widespread application. In related technologies, the training samples for these models rely primarily on massive amounts of manual annotation, resulting in low model training efficiency. While some solutions can automatically generate training samples, they often struggle to generate samples that specifically address the quality defects found in multimedia resources, such as images, generated by these models. This poor training sample quality leads to poor model training results, which in turn leads to issues such as poor denoising performance and poor multimedia resource generation. Summary of the Invention
[0003] The present disclosure provides a multimedia resource generation model training and multimedia resource generation method to at least address the technical issues in related technologies such as poor training sample quality, resulting in poor model training results, and thus poor denoising performance of the multimedia resource generation model and poor multimedia resource generation results. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a method for training a multimedia resource generation model is provided, comprising:
[0005] Acquire multiple resource generation description information, preset guidance parameters corresponding to a preset generation model, and scaled guidance parameters corresponding to the preset guidance parameters;
[0006] Inputting each resource generation description information and preset empty description information into the preset generation model, and guiding the preset generation model based on the preset guidance parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0007] Inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information, and at least one resource quality defect corresponding to any one of the negative sample multimedia resources;
[0008] Based on the at least one negative sample multimedia resource corresponding to each resource generation description information and the at least one positive sample multimedia resource corresponding to each resource generation description information, construct at least one positive and negative sample resource pair corresponding to each resource generation description information;
[0009] Based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information, the multimedia resource generation model to be trained is iteratively trained to obtain a multimedia resource generation model.
[0010] In an optional embodiment, the method further includes:
[0011] Obtaining a preset denoising time step corresponding to the preset generation model and a reduced denoising time step corresponding to the preset denoising time step;
[0012] The step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the preset guidance parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one positive sample multimedia resource corresponding to each resource generation description information includes:
[0013] Inputting each resource generation description information and the preset empty description information into the preset generation model, and in the process of performing multimedia resource generation based on the preset denoising time step, fusing the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guiding parameters to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0014] The step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information to obtain at least one negative sample multimedia resource corresponding to each resource generation description information includes:
[0015] Each resource generation description information and the preset empty description information are input into the preset generation model, and in the process of performing multimedia resource generation based on the reduced denoising time step, the corresponding generation result of each resource generation description information and the corresponding generation result of the preset empty description information are fused based on the scaled guidance parameters to obtain at least one negative sample multimedia resource corresponding to each resource generation description information.
[0016] In an optional embodiment, the scaled guidance parameter includes at least one of a reduced guidance parameter or an enlarged guidance parameter; inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information includes:
[0017] Inputting each resource generation description information and the preset empty description information into the preset generation model, and in a multimedia resource generation process, fusing a generation result corresponding to each resource generation description information and a generation result corresponding to the preset empty description information based on the reduced guidance parameter to obtain a first type of negative sample multimedia resource corresponding to each resource generation description information;
[0018] and / or,
[0019] Inputting each resource generation description information and the preset empty description information into the preset generation model, and in a multimedia resource generation process, fusing a generation result corresponding to each resource generation description information and a generation result corresponding to the preset empty description information based on the amplified guidance parameter to obtain a second type of negative sample multimedia resource corresponding to each resource generation description information;
[0020] The at least one negative sample multimedia resource includes the first type of negative sample multimedia resource and / or the second type of negative sample multimedia resource.
[0021] In an optional embodiment, the iterative training of the multimedia resource generation model to be trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model includes:
[0022] Determine a current positive and negative sample pair from the positive and negative sample resource pairs corresponding to the multiple resource generation description information;
[0023] Determine the current noise adding time step from the preset noise adding time step interval;
[0024] Based on the preset noise information corresponding to the current noise addition time step, the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair are respectively subjected to noise addition processing to obtain a noisy positive and negative sample feature pair;
[0025] Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into the multimedia resource generation model to be trained for denoising, thereby obtaining first noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0026] Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into a reference model corresponding to the multimedia resource generation model to be trained for denoising, to obtain second noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0027] Inputting the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained;
[0028] Based on the noise prediction loss information, the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model.
[0029] In an optional embodiment, the preset loss function is a loss function including a temperature coefficient, the temperature coefficient is used to control the degree of preservation of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is positively correlated with the preservation degree; the current denoising time step corresponding to different training rounds during the iterative training of the multimedia resource generation model to be trained is different, and the multimedia resource generation model to be trained is iteratively trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model further includes:
[0030] Determining a current temperature coefficient corresponding to the current noise addition time step from a preset temperature coefficient interval, wherein the current temperature coefficient is negatively correlated with the current noise addition time step;
[0031] Inputting the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained includes:
[0032] The first noise prediction information, the second noise prediction information, and the preset noise information are input into the preset loss function with the current temperature coefficient as the temperature coefficient to obtain the noise prediction loss information.
[0033] In an optional embodiment, the preset generation model is the multimedia resource generation model to be trained.
[0034] In an optional embodiment, the constructing, based on the at least one negative sample multimedia resource corresponding to each resource generation description information and the at least one positive sample multimedia resource corresponding to each resource generation description information, at least one positive and negative sample resource pair corresponding to each resource generation description information includes:
[0035] Each negative sample multimedia resource in the at least one negative sample multimedia resource corresponding to each resource generation description information is combined with each positive sample multimedia resource in the at least one positive sample multimedia resource corresponding to each resource generation description information to obtain at least one positive and negative sample resource pair corresponding to each resource generation description information.
[0036] According to a second aspect of an embodiment of the present disclosure, a method for generating multimedia resources is provided, including:
[0037] Obtain target resource generation description information and preset noise-added multimedia resources;
[0038] The target resource generation description information and the preset noisy multimedia resource are input into a multimedia resource generation model obtained based on the training method of any multimedia resource generation model provided in the first aspect to perform multimedia resource generation processing to obtain the target multimedia resource corresponding to the target resource generation description information.
[0039] According to a third aspect of an embodiment of the present disclosure, there is provided a training device for a multimedia resource generation model, comprising:
[0040] A first data acquisition module is configured to execute acquisition of a plurality of resource generation description information, preset guidance parameters corresponding to a preset generation model, and scaled guidance parameters corresponding to the preset guidance parameters;
[0041] The positive sample multimedia resource generation module is configured to input each resource generation description information and preset empty description information into the preset generation model, and guide the preset generation model to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guidance parameters, so as to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0042] A negative sample multimedia resource generation module is configured to input each resource generation description information and the preset empty description information into the preset generation model, and guide the preset generation model to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the scaled guidance parameter, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information, and at least one resource quality defect corresponding to any negative sample multimedia resource;
[0043] A positive-negative sample pair construction module is configured to execute, based on at least one negative sample multimedia resource corresponding to each resource generation description information and at least one positive sample multimedia resource corresponding to each resource generation description information, to construct at least one positive-negative sample resource pair corresponding to each resource generation description information;
[0044] The model iteration training module is configured to execute iterative training on the multimedia resource generation model to be trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model.
[0045] In an optional embodiment, the device further comprises:
[0046] A denoising time step acquisition module is configured to acquire a preset denoising time step corresponding to the preset generation model and a reduced denoising time step corresponding to the preset denoising time step;
[0047] The positive sample multimedia resource generation module is further configured to input each resource generation description information and the preset empty description information into the preset generation model, and in the process of performing multimedia resource generation based on the preset denoising time step, fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guidance parameters to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0048] The negative sample multimedia resource generation module is also configured to input each resource generation description information and the preset empty description information into the preset generation model, and in the process of performing multimedia resource generation based on the reduced denoising time step, fuse the corresponding generation results of each resource generation description information and the corresponding generation results of the preset empty description information based on the scaled guidance parameters to obtain at least one negative sample multimedia resource corresponding to each resource generation description information.
[0049] In an optional embodiment, the scaled guide parameter includes at least any one of a reduced guide parameter or an enlarged guide parameter; and the negative sample multimedia resource generation module includes:
[0050] A first negative sample generation unit is configured to input each resource generation description information and the preset empty description information into the preset generation model, and, during a multimedia resource generation process, fuse a generation result corresponding to each resource generation description information with a generation result corresponding to the preset empty description information based on the reduced guidance parameter to obtain a first type of negative sample multimedia resource corresponding to each resource generation description information;
[0051] and / or,
[0052] A second negative sample generating unit is configured to input each resource generation description information and the preset empty description information into the preset generation model, and, during a multimedia resource generation process, fuse a generation result corresponding to each resource generation description information with a generation result corresponding to the preset empty description information based on the amplified guidance parameter to obtain a second type of negative sample multimedia resource corresponding to each resource generation description information;
[0053] The at least one negative sample multimedia resource includes the first type of negative sample multimedia resource and / or the second type of negative sample multimedia resource.
[0054] In an optional embodiment, the model iterative training module includes:
[0055] a current positive-negative sample pair determining unit, configured to determine a current positive-negative sample pair from the positive-negative sample resource pairs corresponding to the description information generated from the plurality of resources;
[0056] A current noise addition time step determining unit is configured to determine a current noise addition time step from a preset noise addition time step interval;
[0057] a noise processing unit configured to perform noise processing on the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair based on preset noise information corresponding to the current noise adding time step, thereby obtaining a noisy positive and negative sample feature pair;
[0058] A first denoising processing unit is configured to input the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into the multimedia resource generation model to be trained for denoising, thereby obtaining first noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0059] A second denoising processing unit is configured to perform denoising processing on the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs, inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the reference model corresponding to the multimedia resource generation model to be trained, to obtain second noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0060] a noise prediction loss information determining unit, configured to input the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained;
[0061] The iterative training unit is configured to perform iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information to obtain the multimedia resource generation model.
[0062] In an optional embodiment, the preset loss function is a loss function including a temperature coefficient, the temperature coefficient is used to control the degree of preservation of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is positively correlated with the preservation degree; during the iterative training process of the multimedia resource generation model to be trained, different training rounds correspond to different current denoising time steps, and the model iterative training module further includes:
[0063] a current temperature coefficient determination module, configured to determine a current temperature coefficient corresponding to the current noise addition time step from a preset temperature coefficient interval, wherein the current temperature coefficient is negatively correlated with the current noise addition time step;
[0064] The noise prediction loss information determining unit is specifically configured to input the first noise prediction information, the second noise prediction information, and the preset noise information into the preset loss function with the current temperature coefficient as the temperature coefficient to obtain the noise prediction loss information.
[0065] In an optional embodiment, the preset generation model is the multimedia resource generation model to be trained.
[0066] In an optional embodiment, the positive-negative sample pair construction module is specifically configured to perform pairwise combinations of each negative sample multimedia resource in at least one negative sample multimedia resource corresponding to each resource generation description information and each positive sample multimedia resource in at least one positive sample multimedia resource corresponding to each resource generation description information, to obtain at least one positive-negative sample resource pair corresponding to each resource generation description information.
[0067] According to a fourth aspect of an embodiment of the present disclosure, there is provided a multimedia resource generating apparatus, comprising:
[0068] The second data acquisition module is configured to acquire target resources, generate description information, and preset noise-added multimedia resources;
[0069] The multimedia resource generation module is configured to execute multimedia resource generation processing by inputting the target resource generation description information and the preset noisy multimedia resource into a multimedia resource generation model obtained by the training method of any multimedia resource generation model provided by the first aspect, and obtain the target multimedia resource corresponding to the target resource generation description information.
[0070] According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a method as described in any one of the first or second aspects above.
[0071] According to the sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is capable of executing any one of the methods in the first aspect or the second aspect of the embodiment of the present disclosure.
[0072] According to a seventh aspect of an embodiment of the present disclosure, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the method as described in any one of the first or second aspects above.
[0073] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0074] During the iterative training process of the multimedia resource generation model, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the preset guiding parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, so as to obtain at least one positive sample multimedia resource corresponding to each resource generation description information. Then, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the scaled guiding parameters corresponding to the preset guiding parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, so as to obtain at least one negative sample corresponding to each resource generation description information. Multimedia resources, and any negative sample multimedia resource corresponds to at least one resource quality defect, and positive and negative samples covering different quality defects can be automatically generated, which greatly improves the model training efficiency. At the same time, negative sample multimedia resources with different resource quality defects can be combined to ensure that at least one negative sample multimedia resource and at least one positive sample multimedia resource corresponding to the description information generated based on each resource are guaranteed. The positive and negative sample resource pairs constructed can be used to conduct more targeted model training in the iterative training process of the multimedia resource generation model to be trained, and to solve various resource quality defect problems of resources generated by generative models such as diffusion models, thereby improving the model training effect, and thus improving the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.
[0075] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0077] Figure 1is a schematic diagram showing an application environment according to an exemplary embodiment;
[0078] Figure 2 is a flowchart of a method for training a multimedia resource generation model according to an exemplary embodiment;
[0079] Figure 3 This is a flowchart illustrating a method of iteratively training a multimedia resource generation model to be trained based on positive and negative sample resource pairs corresponding to multiple resource generation description information to obtain a multimedia resource generation model according to an exemplary embodiment;
[0080] Figure 4 This is a flowchart illustrating another method of iteratively training a multimedia resource generation model to be trained based on positive and negative sample resource pairs corresponding to multiple resource generation description information to obtain a multimedia resource generation model according to an exemplary embodiment;
[0081] Figure 5 is a flowchart of another method for training a multimedia resource generation model provided according to an exemplary embodiment;
[0082] Figure 6 is a flowchart of a method for generating multimedia resources according to an exemplary embodiment;
[0083] Figure 7 is a block diagram of a training device for a multimedia resource generation model according to an exemplary embodiment;
[0084] Figure 8 is a block diagram of a multimedia resource generation device according to an exemplary embodiment;
[0085] Figure 9 is a block diagram of an electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment;
[0086] Figure 10 It is a block diagram of another electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. DETAILED DESCRIPTION
[0087] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0088] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0089] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0090] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an application environment according to an exemplary embodiment. The application environment may include a terminal 100 and a server 200 .
[0091] In an optional embodiment, the terminal 100 can be used to provide multimedia resource generation services to any user. Specifically, the terminal 100 may include, but is not limited to, electronic devices such as smartphones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on these electronic devices, such as applications. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, Windows, etc.
[0092] In an optional embodiment, the server 200 can provide background services for the terminal 100. Specifically, the server 200 can be used to pre-train a multimedia resource generation model and provide multimedia resource generation services to the terminal 100 based on the trained multimedia resource generation model. Specifically, the server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0093] In addition, it should be noted that Figure 1 What is shown is only one application environment provided by the present disclosure. In actual application, other application environments may also be included.
[0094] In the embodiments of this specification, the terminal 100 and the server 200 may be directly or indirectly connected via wired or wireless communication, which is not limited in this disclosure.
[0095] Figure 2 This is a flowchart of a method for training a multimedia resource generation model according to an exemplary embodiment. The method can be applied to terminals, servers, etc. Figure 2 As shown, the method may include the following steps:
[0096] In step S201 , a plurality of resource generation description information, preset guiding parameters corresponding to a preset generation model, and scaled guiding parameters corresponding to the preset guiding parameters are obtained.
[0097] In a specific embodiment, the resource generation description information may be description information of a multimedia resource to be generated. Specifically, a multimedia resource is at least one resource capable of constituting multimedia. Exemplarily, multimedia resources may include static resources such as images, or dynamic resources such as videos, i.e., they may include at least one of images or videos. Optionally, the resource generation description information may include text, images, etc.
[0098] In a specific embodiment, the preset generation model may be a generative model for generating positive and negative sample multimedia resources; the specific model structure may be set in combination with actual applications. For example, the preset generation model may be a preset diffusion model (DM).
[0099] In an optional embodiment, the above-mentioned preset generation model may be a generation model of the multimedia resource to be trained.
[0100] In the above embodiment, the multimedia resource generation model to be trained is used as a preset generation model for generating positive and negative sample multimedia resources, so that the generated samples can include resource quality defects that may exist in the multimedia resource generation model to be trained itself, thereby generating training samples more specifically and better improving the controllability of sample quality.
[0101] In a specific embodiment, preset guidance parameters can be used to guide a preset generation model to generate a multimedia resource that conforms to any resource generation description information and preset empty description information (i.e., a multimedia resource that conforms to the description of the resource generation description information). Specifically, the preset guidance parameters can guide the preset generation model to generate a generation result (multimedia resource) that conforms to the resource generation description information by controlling the weight of the generation result of the preset generation model for any resource generation description information and the generation result for the preset empty description information. Specifically, the preset empty description information can be empty description information, such as empty text.
[0102] In a specific embodiment, the preset guidance parameters can be determined by debugging the guidance parameters of the preset generation model. Specifically, the guidance parameters of the preset generation model can be continuously adjusted, and the multimedia resource generation process can be performed in combination with the adjusted preset generation model, and the corresponding guidance parameters when the generated multimedia resource meets the preset quality conditions can be used as the preset guidance parameters; specifically, the multimedia resource that meets the preset quality conditions can be a multimedia resource that meets the description of the corresponding resource generation description information. Specifically, the preset quality conditions can be set in combination with actual applications. For example, the preset quality conditions can include at least any one of a preset lower limit threshold of the correlation between the multimedia resource and the corresponding resource generation description information, a preset lower limit threshold of the user's satisfaction with the multimedia resource, and a preset lower limit threshold of the quality data of the resource visual attribute; optionally, the above-mentioned correlation can be pre-set by relevant personnel, or can be obtained by extracting the respective features (feature vectors) of the multimedia resource and the corresponding resource generation description information, and using the distance between the features as the corresponding correlation; specifically, the distance between the features can include but is not limited to Euclidean distance, Manhattan distance, etc. Specifically, the user's overall satisfaction with a multimedia resource can be pre-set by relevant personnel. Specifically, the resource visual attributes can be attributes that reflect the visual characteristics of the resource. For example, if the multimedia resource is an image, the resource visual attributes can include at least one of image clarity and image saturation; if the multimedia resource is a video, the resource visual attributes can include at least one of video frame clarity, video frame saturation, and smoothness between video frames. Optionally, the quality data of the resource visual attributes can be pre-set by relevant personnel or can be identified using artificial intelligence technology.
[0103] In a specific embodiment, the scaled guidance parameter corresponding to the preset guidance parameter may include at least any one of the reduced guidance parameter and the enlarged guidance parameter; the reduced guidance parameter may be the guidance parameter after the preset guidance parameter is reduced, and the enlarged guidance parameter may be the guidance parameter after the preset guidance parameter is enlarged.
[0104] In step S203, each resource generation description information and preset empty description information are input into the preset generation model, and the preset generation model is guided based on the preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information to obtain at least one positive sample multimedia resource corresponding to each resource generation description information.
[0105] In a specific embodiment, one or more positive sample multimedia resources (at least one positive sample multimedia resource) can be generated for each resource generated description information. In a specific embodiment, based on preset guidance parameters, a preset generation model is guided to fuse the generation results corresponding to each resource generated description information with the generation results corresponding to the preset empty description information. The at least one positive sample multimedia resource corresponding to each resource generated description information can be obtained by combining the following formula:
[0106]
[0107] in, A positive sample multimedia resource corresponding to the description information of a resource can be generated. The result can be generated for the preset empty description information (i.e., the multimedia resources generated by the preset generation model based on the preset empty description information). The generation result corresponding to the description information of the above-mentioned resource may be generated (ie, the multimedia resource generated by the preset generation model based on the description information of the resource); w may be a preset guiding parameter.
[0108] In step S205, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, thereby obtaining at least one negative sample multimedia resource corresponding to each resource generation description information;
[0109] In a specific embodiment, at least one negative sample multimedia resource may be one or more negative sample multimedia resources; different types of negative sample multimedia resources correspond to different resource quality defects; any one type of negative sample multimedia resource corresponds to at least one resource quality defect, specifically, the resource quality defect may characterize the incompleteness or abnormality of the multimedia resource in visual performance or information transmission; exemplarily, taking the multimedia resource as an image, at least one resource quality defect may include at least any one of image overexposure, image oversaturation, image undersaturation, and image content missing; in the case where the multimedia resource is a video, at least one resource quality defect may include at least any one of video frame image overexposure, video frame image oversaturation, video frame image undersaturation, video frame image content missing, and video frame image non-smoothness.
[0110] In an optional embodiment, the scaled guidance parameter includes at least one of a reduced guidance parameter or an enlarged guidance parameter; accordingly, the step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information may include:
[0111] Inputting each resource generation description information and preset empty description information into a preset generation model, and in a multimedia resource generation process, fusing the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the reduced guidance parameter to obtain a first type of negative sample multimedia resource corresponding to each resource generation description information;
[0112] and / or,
[0113] Inputting each resource generation description information and preset empty description information into a preset generation model, and in the multimedia resource generation process, fusing the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the amplified guidance parameter to obtain a second type of negative sample multimedia resource corresponding to each resource generation description information;
[0114] In a specific embodiment, the at least one negative sample multimedia resource includes a first type of negative sample multimedia resource and / or a second type of negative sample multimedia resource. In an optional embodiment, taking images as an example, resource quality defects corresponding to the first type of negative sample multimedia resource may include missing image content and low image saturation; resource quality defects corresponding to the second type of negative sample multimedia resource may include overexposure and oversaturation.
[0115] In a specific embodiment, the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information are fused based on the reduced guidance parameters to obtain the first type of negative sample multimedia resource corresponding to each resource generation description information, and the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information are fused based on the amplified guidance parameters to obtain the specific refinement of the second type of negative sample multimedia resource corresponding to each resource generation description information. Please refer to the above-mentioned specific refinement of the fusion of the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guidance parameters to guide the preset generation model to obtain at least one positive sample multimedia resource corresponding to each resource generation description information, which will not be repeated here.
[0116] In the above embodiment, combined with at least any one of the guidance parameters after reduction or the guidance parameters after amplification, the preset generation model is guided to fuse the generation results corresponding to the resource generation description information and the generation results corresponding to the preset empty description information. While realizing the automatic generation of negative samples, the preset generation model can generate negative sample multimedia resources with different resource quality defects, thereby enabling more targeted model training to ensure the multimedia resource generation effect of the trained multimedia resource generation model.
[0117] In step S207, based on at least one negative sample multimedia resource corresponding to each resource generation description information and at least one positive sample multimedia resource corresponding to each resource generation description information, at least one positive and negative sample resource pair corresponding to each resource generation description information is constructed;
[0118] In a specific embodiment, any positive and negative sample resource pair corresponding to each resource generation description information may include a positive sample multimedia resource and a negative sample multimedia resource corresponding to the resource generation description information.
[0119] In an optional embodiment, based on the at least one negative sample multimedia resource corresponding to each resource generation description information and the at least one positive sample multimedia resource corresponding to each resource generation description information, constructing at least one positive and negative sample resource pair corresponding to each resource generation description information may include:
[0120] Each negative sample multimedia resource in the at least one negative sample multimedia resource corresponding to each resource generation description information is combined with each positive sample multimedia resource in the at least one positive sample multimedia resource corresponding to each resource generation description information to obtain at least one positive and negative sample resource pair corresponding to each resource generation description information.
[0121] In the above embodiment, each negative sample multimedia resource in the at least one negative sample multimedia resource corresponding to each resource generation description information is combined with each positive sample multimedia resource in the at least one positive sample multimedia resource corresponding to each resource generation description information to obtain at least one positive and negative sample resource pair corresponding to each resource generation description information. This can ensure that the positive and negative sample resource pairs used for multimedia resource generation model training more comprehensively cover various resource quality defects, and thus the model training can be more targeted to ensure the multimedia resource generation effect of the trained multimedia resource generation model.
[0122] In step S209 , the multimedia resource generation model to be trained is iteratively trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain a multimedia resource generation model.
[0123] In an optional embodiment, the multimedia resource generation model to be trained may be a generative model to be trained. The specific model structure may be set in combination with actual applications. For example, the multimedia resource generation model to be trained may be a diffusion model (DM) to be trained.
[0124] In an optional embodiment, if Figure 3 As shown, the above-mentioned positive and negative sample resource pairs corresponding to the multiple resource generation description information are iteratively trained on the multimedia resource generation model to be trained, and the multimedia resource generation model obtained may include:
[0125] S301: Determine a current positive and negative sample pair from the positive and negative sample resource pairs corresponding to the description information generated from multiple resources;
[0126] S303: Determine the current noise adding time step from the preset noise adding time step interval;
[0127] S305: Based on the preset noise information corresponding to the current noise addition time step, the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair are respectively subjected to noise addition processing to obtain a noisy positive and negative sample feature pair;
[0128] S307: Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into the multimedia resource generation model to be trained for denoising, thereby obtaining first noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0129] S309: Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into a reference model corresponding to the multimedia resource generation model to be trained for denoising, thereby obtaining second noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0130] S311: Inputting the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained;
[0131] S313: Based on the noise prediction loss information, iteratively train the multimedia resource generation model to be trained to obtain a multimedia resource generation model.
[0132] In a specific embodiment, the current positive and negative sample pairs may be the positive and negative sample resource pairs corresponding to the current training round during the iterative training (multi-round training) of the multimedia resource generation model. Specifically, the current positive and negative sample pairs may be randomly determined from the positive and negative sample resource pairs corresponding to the multiple resource generation description information, or may be determined from the positive and negative sample resource pairs corresponding to the multiple resource generation description information according to a preset sample pair selection rule. For example, positive and negative sample pairs that have not participated in training may be sequentially selected from the positive and negative sample resource pairs corresponding to the multiple resource generation description information as the current positive and negative sample pairs. Specifically, the positive and negative sample resource pairs corresponding to the multiple resource generation description information may be a collection of positive and negative sample resource pairs used to train the multimedia resource generation model. Specifically, the current positive and negative sample pairs may include at least one positive and negative sample resource pair.
[0133] In a specific embodiment, the preset noise addition time step interval can be a selection interval of the noise addition time step during the iterative training of the multimedia resource generation model to be trained; specifically, the current noise addition time step can be the noise addition time step of the current training round, and specifically, the noise addition time step can represent the intensity of the noise addition. Optionally, according to a preset sampling interval, during the iterative training of the multimedia resource generation model to be trained, the current noise addition time step corresponding to each training round can be selected from the preset noise addition time step interval in sequence, increasing or decreasing. The current noise addition time step corresponding to each training round can also be randomly selected, but it is necessary to ensure that different noise addition time steps are selected from the preset noise addition time step interval during the iterative training process.
[0134] In a specific embodiment, the preset noise information may be noise information of the noise intensity corresponding to the current noise addition time step, for example, Gaussian noise information.
[0135] In a specific embodiment, the noise addition process is often performed in the latent space. Accordingly, the above-mentioned noise addition process is performed on the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair respectively based on the preset noise information corresponding to the current noise addition time step, and the noisy positive and negative sample feature pairs are obtained, which may include: based on the preset noise information corresponding to the current noise addition time step, the original latent space features corresponding to the positive sample multimedia resource and the negative sample multimedia resource are respectively noised to obtain the corresponding noisy positive and negative sample feature pairs. Specifically, the original latent space features corresponding to any multimedia resource may be the resource features after encoding the multimedia resource. Any noisy positive and negative sample feature pair may include the features of the original latent space features of the corresponding positive sample multimedia resource after being noised based on the preset noise information and the features of the original latent space features of the corresponding negative sample multimedia resource after being noised based on the preset noise information.
[0136] In a specific embodiment, the reference model corresponding to the multimedia resource generation model to be trained can be the multimedia resource generation model to be trained in the initial state (that is, the multimedia resource generation model to be trained that has not been iteratively updated), and the reference model remains unchanged during the iterative training process of the multimedia resource generation model to be trained.
[0137] In a specific embodiment, the process of generating multimedia resources by the multimedia resource generation model to be trained (i.e., the process of performing denoising processing) may include a process of predicting noise information and a process of removing noise information based on the predicted noise information; during the training phase of the multimedia resource generation model to be trained, only the predicted noise information (i.e., noise prediction information) may be output.
[0138] In a specific embodiment, the first noise prediction information may be noise information predicted by the multimedia resource generation model to be trained; the second noise prediction information may be noise information predicted by a reference model corresponding to the multimedia resource generation model to be trained.
[0139] In a specific embodiment, the preset loss function can be configured based on actual applications, for example, a mean square error loss function. Specifically, the first noise prediction information, the second noise prediction information, and the preset noise information can be input into the preset loss function to obtain noise prediction loss information. Specifically, the noise prediction loss information can represent the noise prediction performance (i.e., denoising performance) of the current multimedia resource to be trained.
[0140] In an optional embodiment, the preset loss function is a loss function including a temperature coefficient, and the temperature coefficient is used to control the degree of preservation of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is positively correlated with the preservation degree; during the iterative training process of the multimedia resource generation model to be trained, the current denoising time step corresponding to different training rounds is different, such as Figure 4 As shown, based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information, iteratively training the multimedia resource generation model to be trained to obtain the multimedia resource generation model may also include:
[0141] S315: Determine the current temperature coefficient corresponding to the current noise addition time step from the preset temperature coefficient range;
[0142] Accordingly, the inputting of the first noise prediction information, the second noise prediction information, and the preset noise information into the preset loss function to obtain the noise prediction loss information corresponding to the multimedia resource generation model to be trained may include:
[0143] The first noise prediction information, the second noise prediction information, and the preset noise information are input into a preset loss function with a current temperature coefficient as a temperature coefficient to obtain noise prediction loss information.
[0144] In a specific embodiment, the preset temperature coefficient interval can be a preset interval of temperature coefficients selected during the iterative training of the multimedia resource generation model to be trained; the current temperature coefficient is negatively correlated with the current noise addition time step. Specifically, each noise addition time step in the preset noise addition time step interval corresponds to a temperature coefficient in the preset temperature coefficient interval; specifically, the time step lower limit value of the preset noise addition time step interval corresponds to the temperature coefficient upper limit value of the preset temperature coefficient interval; and the time step upper limit value of the preset noise addition time step interval corresponds to the temperature coefficient lower limit value of the preset temperature coefficient interval.
[0145] In a specific embodiment, the loss function including the temperature coefficient can be set in combination with the actual application. For example, the loss function including the temperature coefficient (preset loss function) is as follows:
[0146]
[0147] in, is the noise prediction loss information when the noise adding time step is t; Represents the sigmoid function; represents the temperature coefficient; Indicates the upper limit of the preset noise adding time step interval; and is the preset hyperparameter; Represents the noise information corresponding to the positive sample multimedia resource (the preset noise information corresponding to the noise addition time step t); Indicates that the multimedia resource generation model to be trained is based on the positive sample multimedia resource when the noise time step is t Predicted noise information (first preset noise information); Indicates that the reference model is based on the positive sample multimedia resources when the noise time step is t predicted noise information (second preset noise information); Represents the noise information corresponding to the negative sample multimedia resource (preset noise information); Indicates that the multimedia resource generation model to be trained is based on the negative sample multimedia resource when the noise time step is t Predicted noise information (first preset noise information); Indicates that the reference model is based on negative sample multimedia resources when the noise time step is t Predicted noise information (second preset noise information).
[0148] In the above embodiment, a loss function including a temperature coefficient is used as a preset loss function in the iterative training process of the multimedia resource generation model to be trained, and the temperature coefficient is used to control the degree of retention of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is negatively correlated with the noise addition time step. The linear dynamic relationship between the temperature coefficient and the noise addition time step can be combined to set different temperature coefficients for different time steps. This can ensure that there are more possibilities for model updates when the noise is high (when the noise addition intensity is large), and when the noise is low (when the noise addition intensity is small), the original model performance is maintained as much as possible. On the basis of ensuring the original performance of the model, the model can also learn diverse features, thereby better improving the training effect of the model and the multimedia resource generation effect.
[0149] In a specific embodiment, based on the noise prediction loss information, the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model, which may include: updating the model parameters of the multimedia resource generation model to be trained according to the noise prediction loss information, and performing the next training round based on the multimedia resource generation model to be trained after the updated model parameters, that is, repeating the above steps S301-step S311, and the step of updating the model parameters of the multimedia resource generation model to be trained according to the noise prediction loss information, until the preset convergence condition is met, and using the multimedia resource generation model to be trained when the preset convergence condition is met as the multimedia resource generation model.
[0150] In a specific embodiment, the preset convergence conditions can be set in combination with actual applications, for example, the noise prediction loss information is less than a first preset threshold, or the difference between the pairwise noise prediction loss information obtained in the most recent preset round is less than a second preset threshold, or the training round reaches a third preset threshold, etc.
[0151] In the above embodiment, during the iterative training of the multimedia resource generation model, the current positive and negative sample pairs are determined from the positive and negative sample resource pairs corresponding to the multiple resource generation description information; the current noise addition time step is determined from the preset noise addition time step interval; based on the preset noise information corresponding to the current noise addition time step, the positive sample multimedia resources and the negative sample multimedia resources in the current positive and negative sample pairs are noised respectively to obtain the noisy positive and negative sample feature pairs, and the resource generation description information corresponding to the noisy positive and negative sample feature pairs and the noisy positive and negative sample feature pairs are input into the multimedia resource generation model to be trained for denoising processing to obtain the first noise prediction information corresponding to the noisy positive and negative sample feature pairs, and the noisy positive and negative sample feature pairs are input into the multimedia resource generation model to be trained for denoising processing to obtain the first noise prediction information corresponding to the noisy positive and negative sample feature pairs. The resource generation description information corresponding to the sample feature pairs and the noisy positive and negative sample feature pairs is input into the reference model corresponding to the multimedia resource generation model to be trained for denoising processing to obtain the second noise prediction information corresponding to the noisy positive and negative sample feature pairs, and the first noise prediction information, the second noise prediction information and the preset noise information are input into the preset loss function to obtain noise prediction loss information, and based on the noise prediction loss information, the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model that can be combined with the noise information predicted by the reference model to guide the multimedia resource generation model to be trained to retain the original performance during denoising, and also allow the model to learn diverse features.
[0152] In an optional embodiment, if Figure 5 As shown, the above method may further include:
[0153] S211: Obtaining a preset denoising time step corresponding to a preset generation model and a reduced denoising time step corresponding to the preset denoising time step;
[0154] Accordingly, the above-mentioned inputting each resource generation description information and preset empty description information into the preset generation model, and guiding the preset generation model based on the preset guidance parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one positive sample multimedia resource corresponding to each resource generation description information, including:
[0155] Inputting each resource generation description information and preset empty description information into a preset generation model, and in a multimedia resource generation process based on a preset denoising time step, fusing the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on preset guiding parameters to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0156] Accordingly, the above-mentioned inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information, including:
[0157] Each resource generation description information and preset empty description information are input into a preset generation model, and in the process of multimedia resource generation based on the reduced denoising time step, the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information are fused based on the scaled guidance parameters to obtain at least one negative sample multimedia resource corresponding to each resource generation description information.
[0158] In a specific embodiment, the preset denoising time step may be a preset denoising time step that can ensure the multimedia resource generation effect of the preset generation model;
[0159] In a specific embodiment, each resource generation description information and preset empty description information are input into a preset generation model, and in the process of multimedia resource generation based on a preset denoising time step, the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information are fused based on preset guidance parameters to obtain the specific details of at least one positive sample multimedia resource corresponding to each resource generation description information. Please refer to the above-mentioned relevant description and will not be repeated here. That is, in the process of generating the generation results corresponding to the resource generation description information and the generation results corresponding to the preset empty description information, the preset denoising time step is used as the denoising intensity.
[0160] In a specific embodiment, each resource generation description information and preset empty description information are input into a preset generation model, and in the process of multimedia resource generation based on the reduced denoising time step, the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information are fused based on the scaled guide parameters to obtain the specific refinement of at least one negative sample multimedia resource corresponding to each resource generation description information. Please refer to the above-mentioned relevant description and will not be repeated here. That is, in the process of generating the generation result corresponding to the resource generation description information and the generation result corresponding to the preset empty description information, the reduced denoising time step corresponding to the preset denoising time step is used as the denoising intensity.
[0161] In the above embodiment, in the process of generating positive sample multimedia resources, combining the preset guiding parameters and the higher preset denoising time step can better guarantee the effect of the generated multimedia resources; in the process of generating negative sample multimedia resources, combining the scaled guiding parameters and the lower denoising time step can better guarantee the coverage of resource quality defects of the generated multimedia resources, better improve the quality of the positive and negative sample resource pairs, and thus better improve the multimedia resource generation effect of the subsequent multimedia resource generation model obtained by training based on the positive and negative sample resource pairs.
[0162] It can be seen from the technical solutions provided by the above embodiments of this specification that in the iterative training process of the multimedia resource generation model in this specification, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, so as to obtain at least one positive sample multimedia resource corresponding to each resource generation description information. Then, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the scaled guidance parameters corresponding to the preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, so as to obtain each resource generation description information. At least one negative sample multimedia resource corresponds to the above information, and any negative sample multimedia resource corresponds to at least one resource quality defect. Positive and negative samples covering different quality defects can be automatically generated, which greatly improves the model training efficiency. At the same time, negative sample multimedia resources with different resource quality defects can be combined to ensure that at least one negative sample multimedia resource and at least one positive sample multimedia resource corresponding to the description information are generated based on each resource. The positive and negative sample resource pairs constructed can be used to conduct more targeted model training in the iterative training process of the multimedia resource generation model to be trained, and to solve various resource quality defect problems of resources generated by generative models such as diffusion models, thereby improving the model training effect, and thus improving the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.
[0163] Based on the training method of the multimedia resource generation model provided in the embodiment of the present application, a multimedia resource generation method provided in the embodiment of the present application is introduced. Figure 6 is a flowchart of a method for generating multimedia resources according to an exemplary embodiment. The method can be applied to a terminal, a server, etc. Figure 6 As shown, the method may include the following steps:
[0164] In step S601, target resource generation description information and preset noise-added multimedia resources are obtained;
[0165] In a specific embodiment, the target resource generation description information may be description information of a multimedia resource to be generated; the target resource generation description information may include text or an image. In a specific embodiment, the preset noisy multimedia resource may be a multimedia resource generated based on preset noise; for example, a multimedia resource such as a Gaussian noise image generated based on Gaussian noise.
[0166] In step S603, the target resource generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model to perform multimedia resource generation processing, and a target multimedia resource corresponding to the target resource generation description information is obtained.
[0167] In a specific embodiment, the target resource generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model for multimedia resource generation processing, and obtaining the target multimedia resource corresponding to the target resource generation description information may include: inputting the target resource generation description information and the preset noisy multimedia resource into the multimedia resource generation model, the multimedia resource generation model performs noise prediction based on the target resource generation description information and the preset noisy multimedia resource to obtain target noise prediction information, and denoising the preset noisy multimedia resource based on the target noise prediction information to obtain the target multimedia resource.
[0168] In addition, it should be noted that the multiple in the embodiments of the present application can be at least two, and the multiple can be at least two.
[0169] It can be seen from the technical solutions provided by the above embodiments of this specification that, in the multimedia resource generation process, the target resource generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model for multimedia resource generation processing. During the iterative training of the multimedia resource generation model, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, and at least one positive sample multimedia resource corresponding to each resource generation description information is obtained. Then, each resource generation description information and the preset empty description information are input into the preset generation model, and the preset generation model is guided based on the scaled guidance parameters corresponding to the preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information. By fusion of the generation results corresponding to the preset empty description information, at least one negative sample multimedia resource corresponding to each resource generation description information can be obtained, and any negative sample multimedia resource corresponds to at least one resource quality defect, and positive and negative samples covering different quality defects can be automatically generated, which greatly improves the model training efficiency. At the same time, negative sample multimedia resources with different resource quality defects can be combined to ensure that at least one negative sample multimedia resource and at least one positive sample multimedia resource corresponding to each resource generation description information are constructed. In the iterative training process of the multimedia resource generation model to be trained, the model training can be more targeted, and various resource quality defect problems of resources generated by generative models such as diffusion models can be solved in a targeted manner, thereby improving the model training effect, and then improving the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.
[0170] Figure 7 FIG1 is a block diagram of a training device for a multimedia resource generation model according to an exemplary embodiment. Figure 7 , the device comprises:
[0171] The first data acquisition module 710 is configured to execute acquisition of a plurality of resource generation description information, preset guidance parameters corresponding to a preset generation model, and scaled guidance parameters corresponding to the preset guidance parameters;
[0172] The positive sample multimedia resource generation module 720 is configured to input each resource generation description information and preset empty description information into a preset generation model, and guide the preset generation model based on preset guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, thereby obtaining at least one positive sample multimedia resource corresponding to each resource generation description information;
[0173] The negative sample multimedia resource generation module 730 is configured to input each resource generation description information and preset empty description information into a preset generation model, and guide the preset generation model based on the scaled guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, thereby obtaining at least one negative sample multimedia resource corresponding to each resource generation description information and at least one resource quality defect corresponding to any negative sample multimedia resource;
[0174] The positive-negative sample pair construction module 740 is configured to execute, based on at least one negative sample multimedia resource corresponding to each resource generation description information and at least one positive sample multimedia resource corresponding to each resource generation description information, to construct at least one positive-negative sample resource pair corresponding to each resource generation description information;
[0175] The model iteration training module 750 is configured to execute iterative training on the multimedia resource generation model to be trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model.
[0176] In an optional embodiment, the above device further includes:
[0177] A denoising time step acquisition module is configured to execute acquisition of a preset denoising time step corresponding to a preset generation model and a reduced denoising time step corresponding to the preset denoising time step;
[0178] The positive sample multimedia resource generation module 720 is further configured to input each resource generation description information and the preset empty description information into a preset generation model, and in the process of generating multimedia resources based on the preset denoising time step, fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guidance parameters to obtain at least one positive sample multimedia resource corresponding to each resource generation description information;
[0179] The negative sample multimedia resource generation module 730 is also configured to input each resource generation description information and preset empty description information into a preset generation model, and in the process of multimedia resource generation based on the reduced denoising time step, fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information based on the scaled guidance parameters to obtain at least one negative sample multimedia resource corresponding to each resource generation description information.
[0180] In an optional embodiment, the scaled guide parameter includes at least one of a reduced guide parameter and an enlarged guide parameter; and the negative sample multimedia resource generation module 730 includes:
[0181] The first negative sample generation unit is configured to input each resource generation description information and preset empty description information into a preset generation model, and fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the reduced guidance parameter during the multimedia resource generation process to obtain a first type of negative sample multimedia resource corresponding to each resource generation description information;
[0182] and / or,
[0183] The second negative sample generation unit is configured to input each resource generation description information and the preset empty description information into a preset generation model, and in a multimedia resource generation process, fuse the generation result corresponding to each resource generation description information with the generation result corresponding to the preset empty description information based on the amplified guidance parameter to obtain a second type of negative sample multimedia resource corresponding to each resource generation description information;
[0184] The at least one negative sample multimedia resource includes a first type of negative sample multimedia resource and / or a second type of negative sample multimedia resource.
[0185] In an optional embodiment, the model iterative training module 750 includes:
[0186] A current positive-negative sample pair determination unit is configured to determine a current positive-negative sample pair from the positive-negative sample resource pairs corresponding to the description information generated from the multiple resources;
[0187] A current noise addition time step determining unit is configured to determine a current noise addition time step from a preset noise addition time step interval;
[0188] A noise processing unit is configured to perform noise processing on the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair based on preset noise information corresponding to the current noise adding time step, thereby obtaining a noisy positive and negative sample feature pair;
[0189] A first denoising processing unit is configured to perform denoising processing on the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs, inputting the noisy positive and negative sample feature pairs into the multimedia resource generation model to be trained, and obtain first noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0190] The second denoising processing unit is configured to perform denoising processing on the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs, inputting the noisy positive and negative sample feature pairs into a reference model corresponding to the multimedia resource generation model to be trained, and obtaining second noise prediction information corresponding to the noisy positive and negative sample feature pairs;
[0191] a noise prediction loss information determining unit configured to input the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained;
[0192] The iterative training unit is configured to perform iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information to obtain the multimedia resource generation model.
[0193] In an optional embodiment, the preset loss function is a loss function including a temperature coefficient, the temperature coefficient is used to control the degree of preservation of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is positively correlated with the preservation degree; during the iterative training process of the multimedia resource generation model to be trained, different training rounds correspond to different current denoising time steps, and the model iterative training module 750 further includes:
[0194] A current temperature coefficient determination module is configured to determine a current temperature coefficient corresponding to a current noise addition time step from a preset temperature coefficient interval, wherein the current temperature coefficient is negatively correlated with the current noise addition time step;
[0195] The noise prediction loss information determining unit is specifically configured to input the first noise prediction information, the second noise prediction information and the preset noise information into a preset loss function with a current temperature coefficient as a temperature coefficient to obtain the noise prediction loss information.
[0196] In an optional embodiment, the preset generation model is a multimedia resource generation model to be trained.
[0197] In an optional embodiment, the positive-negative sample pair construction module 740 is specifically configured to perform pairwise combinations of each negative sample multimedia resource in at least one negative sample multimedia resource corresponding to each resource generation description information and each positive sample multimedia resource in at least one positive sample multimedia resource corresponding to each resource generation description information, to obtain at least one positive-negative sample resource pair corresponding to each resource generation description information.
[0198] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0199] Figure 8 FIG. 1 is a block diagram of a multimedia resource generation device according to an exemplary embodiment. Figure 8 , the device comprises:
[0200] The second data acquisition module 810 is configured to acquire target resource generation description information and preset noise-added multimedia resources;
[0201] The multimedia resource generation module 820 is configured to execute multimedia resource generation processing by inputting the target resource generation description information and the preset noisy multimedia resource into the multimedia resource generation model obtained by the training method of any multimedia resource generation model provided above, and obtain the target multimedia resource corresponding to the target resource generation description information.
[0202] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0203] Figure 9 This is a block diagram of an electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 9 As shown. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a training method for a multimedia resource generation model or a multimedia resource generation method is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0204] Figure 10 This is a block diagram of another electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as shown in FIG. Figure 10 As shown. The electronic device includes a processor, a memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a training method for a multimedia resource generation model or a multimedia resource generation method is implemented.
[0205] Those skilled in the art will understand that Figure 9 or Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present disclosure, and does not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0206] In an exemplary embodiment, an electronic device is also provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a training method for a multimedia resource generation model or a multimedia resource generation method as in an embodiment of the present disclosure.
[0207] In an exemplary embodiment, a computer-readable storage medium is also provided. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the training method of the multimedia resource generation model or the multimedia resource generation method in the embodiment of the present disclosure.
[0208] In an exemplary embodiment, a computer program product including instructions is further provided. When the computer program product is run on a computer, the computer is caused to execute the multimedia resource generation model training method or multimedia resource generation method in the embodiments of the present disclosure.
[0209] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, which can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0210] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0211] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for a multimedia resource generation model, characterized in that: include: Acquire multiple resource generation description information, preset guidance parameters corresponding to the preset generation model, and scaled guidance parameters corresponding to the preset guidance parameters; Inputting each resource generation description information and preset empty description information into the preset generation model, and guiding the preset generation model based on the preset guidance parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one positive sample multimedia resource corresponding to each resource generation description information; Inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information, wherein any negative sample multimedia resource corresponds to at least one resource quality defect; Based on the at least one negative sample multimedia resource corresponding to each resource generation description information and the at least one positive sample multimedia resource corresponding to each resource generation description information, construct at least one positive and negative sample resource pair corresponding to each resource generation description information; Based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information, the multimedia resource generation model to be trained is iteratively trained to obtain a multimedia resource generation model.
2. The method for training a multimedia resource generation model according to claim 1, wherein: The method further comprises: Obtaining a preset denoising time step corresponding to the preset generation model and a reduced denoising time step corresponding to the preset denoising time step; The step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the preset guidance parameters to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information, to obtain at least one positive sample multimedia resource corresponding to each resource generation description information includes: Inputting each resource generation description information and the preset empty description information into the preset generation model, and in the process of performing multimedia resource generation based on the preset denoising time step, fusing the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guiding parameters to obtain at least one positive sample multimedia resource corresponding to each resource generation description information; The step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameter to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information to obtain at least one negative sample multimedia resource corresponding to each resource generation description information includes: Each resource generation description information and the preset empty description information are input into the preset generation model, and in the process of performing multimedia resource generation based on the reduced denoising time step, the corresponding generation result of each resource generation description information and the corresponding generation result of the preset empty description information are fused based on the scaled guidance parameters to obtain at least one negative sample multimedia resource corresponding to each resource generation description information.
3. The training method for a multimedia resource generation model according to claim 1 or 2, characterized in that: The scaled guidance parameters include at least one of the reduced guidance parameters and the enlarged guidance parameters; the step of inputting each resource generation description information and the preset empty description information into the preset generation model, and guiding the preset generation model based on the scaled guidance parameters to fuse the generation results corresponding to each resource generation description information and the generation results corresponding to the preset empty description information, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information includes: Inputting each resource generation description information and the preset empty description information into the preset generation model, and in a multimedia resource generation process, fusing a generation result corresponding to each resource generation description information and a generation result corresponding to the preset empty description information based on the reduced guidance parameter to obtain a first type of negative sample multimedia resource corresponding to each resource generation description information; and / or, Inputting each resource generation description information and the preset empty description information into the preset generation model, and in a multimedia resource generation process, fusing a generation result corresponding to each resource generation description information and a generation result corresponding to the preset empty description information based on the amplified guidance parameter to obtain a second type of negative sample multimedia resource corresponding to each resource generation description information; The at least one negative sample multimedia resource includes the first type of negative sample multimedia resource and / or the second type of negative sample multimedia resource.
4. The method for training a multimedia resource generation model according to claim 1 or 2, wherein: The iterative training of the multimedia resource generation model to be trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model includes: Determine a current positive and negative sample pair from the positive and negative sample resource pairs corresponding to the multiple resource generation description information; Determine the current noise adding time step from the preset noise adding time step interval; Based on the preset noise information corresponding to the current noise addition time step, the positive sample multimedia resource and the negative sample multimedia resource in the current positive and negative sample pair are respectively subjected to noise addition processing to obtain a noisy positive and negative sample feature pair; Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into the multimedia resource generation model to be trained for denoising, thereby obtaining first noise prediction information corresponding to the noisy positive and negative sample feature pairs; Inputting the noisy positive and negative sample feature pairs and the resource generation description information corresponding to the noisy positive and negative sample feature pairs into a reference model corresponding to the multimedia resource generation model to be trained for denoising, to obtain second noise prediction information corresponding to the noisy positive and negative sample feature pairs; Inputting the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; Based on the noise prediction loss information, the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model.
5. The method for training a multimedia resource generation model according to claim 4, wherein: The preset loss function is a loss function including a temperature coefficient, the temperature coefficient is used to control the degree of preservation of the original denoising performance of the multimedia resource generation model to be trained, and the temperature coefficient is positively correlated with the preservation degree; the current denoising time steps corresponding to different training rounds during the iterative training of the multimedia resource generation model to be trained are different, and the multimedia resource generation model to be trained is iteratively trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information, to obtain the multimedia resource generation model further comprising: Determining a current temperature coefficient corresponding to the current noise addition time step from a preset temperature coefficient interval, wherein the current temperature coefficient is negatively correlated with the current noise addition time step; Inputting the first noise prediction information, the second noise prediction information, and the preset noise information into a preset loss function to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained includes: The first noise prediction information, the second noise prediction information, and the preset noise information are input into the preset loss function with the current temperature coefficient as the temperature coefficient to obtain the noise prediction loss information.
6. The method for training a multimedia resource generation model according to claim 1 or 2, characterized in that: The preset generation model is the multimedia resource generation model to be trained.
7. The method for training a multimedia resource generation model according to claim 1 or 2, characterized in that: The constructing, based on the at least one negative sample multimedia resource corresponding to each resource generation description information and the at least one positive sample multimedia resource corresponding to each resource generation description information, at least one positive sample resource pair corresponding to each resource generation description information comprises: Each negative sample multimedia resource in the at least one negative sample multimedia resource corresponding to each resource generation description information is combined with each positive sample multimedia resource in the at least one positive sample multimedia resource corresponding to each resource generation description information to obtain at least one positive and negative sample resource pair corresponding to each resource generation description information.
8. A method for generating multimedia resources, characterized in that: include: Obtain target resource generation description information and preset noise-added multimedia resources; The target resource generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model obtained by the training method of the multimedia resource generation model according to any one of claims 1 to 7 to perform multimedia resource generation processing to obtain the target multimedia resource corresponding to the target resource generation description information.
9. A training device for a multimedia resource generation model, characterized in that: include: A first data acquisition module is configured to execute acquisition of a plurality of resource generation description information, preset guidance parameters corresponding to a preset generation model, and scaled guidance parameters corresponding to the preset guidance parameters; The positive sample multimedia resource generation module is configured to input each resource generation description information and preset empty description information into the preset generation model, and guide the preset generation model to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the preset guidance parameters, so as to obtain at least one positive sample multimedia resource corresponding to each resource generation description information; A negative sample multimedia resource generation module is configured to input each resource generation description information and the preset empty description information into the preset generation model, and guide the preset generation model to fuse the generation result corresponding to each resource generation description information and the generation result corresponding to the preset empty description information based on the scaled guidance parameter, to obtain at least one negative sample multimedia resource corresponding to each resource generation description information, wherein any negative sample multimedia resource corresponds to at least one resource quality defect; A positive-negative sample pair construction module is configured to execute, based on at least one negative sample multimedia resource corresponding to each resource generation description information and at least one positive sample multimedia resource corresponding to each resource generation description information, to construct at least one positive-negative sample resource pair corresponding to each resource generation description information; The model iteration training module is configured to execute iterative training on the multimedia resource generation model to be trained based on the positive and negative sample resource pairs corresponding to the multiple resource generation description information to obtain the multimedia resource generation model.
10. A multimedia resource generating device, characterized in that: include: The second data acquisition module is configured to acquire target resources, generate description information, and preset noise-added multimedia resources; The multimedia resource generation module is configured to execute multimedia resource generation processing by inputting the target resource generation description information and the preset noisy multimedia resource into the multimedia resource generation model obtained by the training method of the multimedia resource generation model according to any one of claims 1 to 7, and obtain the target multimedia resource corresponding to the target resource generation description information.
11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of the multimedia resource generation model according to any one of claims 1 to 7 or the multimedia resource generation method according to claim 8.
12. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method for a multimedia resource generation model as described in any one of claims 1 to 7 or the multimedia resource generation method as described in claim 8.
13. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method for training a multimedia resource generation model according to any one of claims 1 to 7 or the method for generating multimedia resources according to claim 8.
Citation Information
Patent Citations
Model training method and device, data processing method and device and electronic equipment
CN114898184A
Image generation method and device based on interaction, electronic equipment and storage medium
CN116306588A