Training method of multimedia resource generation model and multimedia resource generation method

By performing noise addition processing and direct reward prediction in hidden space, and iterative training combined with loss information, the problems of complex and low efficiency of diffusion model training are solved, and efficient denoising and high-quality generation of multimedia resource generation models are achieved.

CN120494015AActive Publication Date: 2025-08-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510983252.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

The existing diffusion model is complex and inefficient in the multimedia resource generation process, resulting in poor generation results and inability to effectively process noise images at different time steps, affecting the model denoising performance and generation effect.

Method used

By performing noise addition processing in hidden space, the characteristics, time steps and generation description information of the noise addition space are directly input into the preset reward model for reward prediction. The multimedia resource generation model is iteratively trained in combination with reward loss information and noise prediction loss information, simplifying the model training steps and improving the model's reward discrimination ability under different noise conditions.

Benefits of technology

The model training process is simplified, the training efficiency and generation effect are improved, and the denoising performance and generation quality of the multimedia resource generation model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494015A_ABST
    Figure CN120494015A_ABST
Patent Text Reader

Abstract

The invention relates to a multimedia resource generation model training method and a multimedia resource generation method, and the training method comprises the steps: carrying out the noise addition of at least one preset multimedia resource corresponding to a plurality of pieces of current generation description information based on the first preset noise information corresponding to a first current noise addition time step; inputting the obtained current noise adding hidden space features, the first current noise adding time step and each piece of current generation description information into a preset reward model for reward prediction to obtain first prediction reward data, and determining reward loss information; determining noise prediction loss information corresponding to the to-be-trained multimedia resource generation model based on the current noise-added hidden space features, each piece of current generation description information and first preset noise information; and performing iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information and the reward loss information to obtain the multimedia resource generation model. According to the embodiment of the invention, the generation effect of the multimedia resource generation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a multimedia resource generation model training and multimedia resource generation method. Background Art

[0002] With the rapid development of artificial intelligence (AI), generative models such as diffusion models have gained widespread application. Diffusion models have become the industry benchmark for generating multimedia resources, such as images. Their core approach is to transform random noise into high-quality images through a gradual denoising process. This process involves forward diffusion (gradually adding Gaussian noise to an image or other multimedia resource) and backward diffusion (gradually removing noise by training a neural network to reconstruct the image or other multimedia resource).

[0003] In related technologies, during diffusion model training, Vision-Language Models (VLMs) are typically used as pixel-level reward models to approximate human preferences. However, when these models are used for human feedback preference optimization (model optimization training), they face the challenge of processing noisy images at different time steps (different noise intensities). Furthermore, they require complex pixel-space transformations, requiring image restoration before reward prediction. This results in complex and inefficient model training. Furthermore, high-noise images cannot be restored to their true form, forcing the model to be updated only with low-noise resources. This results in poor model training results, which in turn leads to poor denoising performance and poor multimedia resource generation results. Summary of the Invention

[0004] The present disclosure provides a multimedia resource generation model training and multimedia resource generation method to at least address the technical issues in related technologies such as the complex model training process, low efficiency, and poor model training results, which in turn lead to poor denoising performance of the multimedia resource generation model and poor multimedia resource generation results. The technical solutions of the present disclosure are as follows: According to a first aspect of an embodiment of the present disclosure, a method for training a multimedia resource generation model is provided, comprising: Obtain at least one preset multimedia resource corresponding to each of the multiple currently generated description information; Noising the at least one preset multimedia resource based on first preset noise information corresponding to the first current noisy time step to obtain at least one current noisy latent space feature corresponding to each currently generated description information; Inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a preset reward model to perform reward prediction, obtaining first predicted reward data, where the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource; Determining reward loss information based on the first predicted reward data; Performing denoising reasoning on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each currently generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; Based on the noise prediction loss information and the reward loss information, the multimedia resource generation model to be trained is iteratively trained to obtain a multimedia resource generation model.

[0005] In an optional embodiment, the preset reward model includes a latent space diffusion model and a reward prediction model; inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the preset reward model for reward prediction to obtain the first predicted reward data includes: Inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the latent space diffusion model to perform resource quality feature extraction processing to obtain a resource quality feature; The resource quality characteristics are input into the reward prediction model to perform reward prediction to obtain the first predicted reward data.

[0006] In an optional embodiment, the method further includes: Acquire at least one sample multimedia resource corresponding to each of the plurality of sample generation description information and preset quality data corresponding to the at least one sample multimedia resource; performing noise processing on the at least one sample multimedia resource based on second preset noise information corresponding to the second current noise addition time step, to obtain at least one sample noisy latent space feature corresponding to each sample generation description information; Inputting the at least one sample noisy latent space feature, the second current noisy time step, and the multiple sample generation description information into the to-be-trained reward model to perform reward prediction, obtaining second predicted reward data, where the second predicted reward data is used to measure the quality of the corresponding sample multimedia resource; determining quality prediction loss information based on the preset quality data and the second prediction reward data; Based on the quality prediction loss information, the reward model to be trained is iteratively trained to obtain the preset reward model.

[0007] In an optional embodiment, when the at least one sample multimedia resource corresponding to each sample generated description information is a plurality of sample multimedia resources, the preset quality data corresponding to the plurality of sample multimedia resources is the quality sorting information corresponding to the plurality of sample multimedia resources, and the quality sorting information is information that is sorted in ascending or descending order according to the quality of the resources corresponding to the plurality of sample multimedia resources.

[0008] In an optional embodiment, determining the quality prediction loss information based on the preset quality data and the second prediction reward data includes: Determining, according to the preset quality data, a positive sample multimedia resource corresponding to each sample generation description information and a negative sample multimedia resource corresponding to each sample generation description information from the multiple sample multimedia resources corresponding to each sample generation description information; The quality prediction loss information is determined according to the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information.

[0009] In an optional embodiment, when each sample generated description information corresponds to at least one sample multimedia resource, the preset quality data corresponding to the at least one sample multimedia resource is data used to measure the quality of each resource of the at least one sample multimedia resource.

[0010] In an optional embodiment, performing denoising inference on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each currently generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained includes: Inputting the at least one current noisy latent space feature and each currently generated description information into the multimedia resource generation model to be trained for denoising, to obtain first noise prediction information; Inputting the at least one current noisy latent space feature and the multiple current generated description information into a reference model corresponding to the multimedia resource generation model to be trained for denoising, to obtain second noise prediction information; The noise prediction loss information is determined according to the first noise prediction information, the second noise prediction information, and the first preset noise information.

[0011] According to a second aspect of an embodiment of the present disclosure, a method for generating multimedia resources is provided, including: Obtain target generation description information and preset noise-added multimedia resources; The target generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model obtained based on the training method of the multimedia resource generation model provided in the first aspect to perform multimedia resource generation processing to obtain the target multimedia resource corresponding to the target generation description information.

[0012] According to a third aspect of an embodiment of the present disclosure, a training device for a multimedia resource generation model is provided, comprising: A multimedia resource acquisition module is configured to acquire at least one preset multimedia resource corresponding to each of the plurality of currently generated description information; a first noise processing module configured to perform noise processing on the at least one preset multimedia resource based on first preset noise information corresponding to a first current noise adding time step, to obtain at least one current noisy latent space feature corresponding to each currently generated description information; A first reward prediction module is configured to perform reward prediction by inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a preset reward model to obtain first predicted reward data, where the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource; a reward loss information determination module, configured to determine reward loss information based on the first predicted reward data; a denoising inference module configured to perform denoising inference on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each currently generated description information, and the first preset noise information, to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; The first model iterative training module is configured to perform iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information and the reward loss information to obtain a multimedia resource generation model.

[0013] In an optional embodiment, the preset reward model includes a latent space diffusion model and a reward prediction model; the first reward prediction module includes: a resource quality feature extraction unit configured to perform resource quality feature extraction processing by inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the latent space diffusion model to obtain a resource quality feature; The reward prediction unit is configured to input the resource quality characteristics into the reward prediction model to perform reward prediction and obtain the first predicted reward data.

[0014] In an optional embodiment, the device further comprises: A training data acquisition module is configured to acquire at least one sample multimedia resource corresponding to each of a plurality of sample generation description information and preset quality data corresponding to the at least one sample multimedia resource; a noise processing module configured to perform noise processing on the at least one sample multimedia resource based on second preset noise information corresponding to a second current noise adding time step, to obtain at least one sample noisy latent space feature corresponding to each sample generated description information; A second reward prediction module is configured to perform reward prediction by inputting the at least one sample noisy latent space feature, the second current noisy time step, and the plurality of sample generation description information into a to-be-trained reward model to obtain second predicted reward data, where the second predicted reward data is used to measure the quality of the corresponding sample multimedia resource; a quality prediction loss information determining module, configured to determine quality prediction loss information based on the preset quality data and the second prediction reward data; The second model iterative training module is configured to perform iterative training on the reward model to be trained based on the quality prediction loss information to obtain the preset reward model.

[0015] In an optional embodiment, when the at least one sample multimedia resource corresponding to each sample generated description information is a plurality of sample multimedia resources, the preset quality data corresponding to the plurality of sample multimedia resources is the quality sorting information corresponding to the plurality of sample multimedia resources, and the quality sorting information is information that is sorted in ascending or descending order according to the quality of the resources corresponding to the plurality of sample multimedia resources.

[0016] In an optional embodiment, the quality prediction loss information determination module includes: a positive and negative sample data acquisition unit configured to determine, based on the preset quality data, a positive sample multimedia resource corresponding to each sample generation description information and a negative sample multimedia resource corresponding to each sample generation description information from the plurality of sample multimedia resources corresponding to each sample generation description information; The quality prediction loss information determination unit is configured to execute the second prediction reward data of the positive sample multimedia resource corresponding to the description information generated by each sample and the second prediction reward data of the negative sample multimedia resource corresponding to the description information generated by each sample to determine the quality prediction loss information.

[0017] In an optional embodiment, when each sample generated description information corresponds to at least one sample multimedia resource, the preset quality data corresponding to the at least one sample multimedia resource is data used to measure the quality of each resource of the at least one sample multimedia resource.

[0018] In an optional embodiment, the denoising inference module includes: A first denoising processing unit is configured to perform denoising processing on the at least one current noisy latent space feature and each currently generated description information as inputs to the multimedia resource generation model to be trained, thereby obtaining first noise prediction information; A second denoising processing unit is configured to perform denoising processing on the at least one current noisy latent space feature and the multiple currently generated description information input into a reference model corresponding to the multimedia resource generation model to be trained, to obtain second noise prediction information; The noise prediction loss information determining unit is configured to determine the noise prediction loss information according to the first noise prediction information, the second noise prediction information and the first preset noise information.

[0019] According to a fourth aspect of an embodiment of the present disclosure, there is provided a multimedia resource generating apparatus, comprising: A data acquisition module is configured to execute acquisition target generation description information and preset noise-added multimedia resources; The multimedia resource generation module is configured to execute multimedia resource generation processing by inputting the target generation description information and the preset noisy multimedia resource into a multimedia resource generation model obtained by the training method of any multimedia resource generation model provided by the first aspect, so as to obtain the target multimedia resource corresponding to the target generation description information.

[0020] According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a method as described in any one of the first or second aspects above.

[0021] According to the sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is capable of executing any one of the methods in the first aspect or the second aspect of the embodiment of the present disclosure. According to a seventh aspect of an embodiment of the present disclosure, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the method as described in any one of the first or second aspects above.

[0022] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects: During the iterative training of the multimedia resource generation model, based on the first preset noise information corresponding to the first current noise addition time step, at least one preset multimedia resource corresponding to each of the multiple currently generated description information is noised in the latent space to obtain at least one current noisy latent space feature corresponding to each currently generated description information, and the first current noisy time step, each currently generated description information and at least one current noisy latent space feature are directly input into the preset reward model for reward prediction. Reward prediction can be performed without restoring multimedia resources such as images, thereby simplifying the model training steps and improving the model training efficiency. In addition, during the iterative training of the model, different noise addition time steps can be selected from the first preset time step interval, which effectively avoids the inability to restore the real image due to high noise latent space features. Effective supervision of high-noise areas can greatly improve the reward discrimination ability of the multimedia resource generation model to be trained under different noise conditions, and determine the reward loss information by combining the first predicted reward data obtained by reward prediction to measure the quality of the corresponding preset multimedia resource; and based on the reward loss information and at least one current noisy latent space feature, each current generated description information and the first preset noise information, perform denoising reasoning on the multimedia resource generation model to be trained, and together with the obtained noise prediction loss information corresponding to the multimedia resource generation model to be trained, iteratively train the multimedia resource generation model to obtain a multimedia resource generation model, which can greatly improve the model training effect, and thereby improve the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.

[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0025] Figure 1 is a schematic diagram showing an application environment according to an exemplary embodiment; Figure 2 is a flowchart of a method for training a multimedia resource generation model according to an exemplary embodiment; Figure 3 is a flowchart illustrating a training preset reward model according to an exemplary embodiment; Figure 4 is a flowchart of a method for generating multimedia resources according to an exemplary embodiment; Figure 5 is a block diagram of a training device for a multimedia resource generation model according to an exemplary embodiment; Figure 6 is a block diagram of a multimedia resource generation device according to an exemplary embodiment; Figure 7 is a block diagram of an electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment; Figure 8 It is a block diagram of another electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. DETAILED DESCRIPTION

[0026] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0027] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0029] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an application environment according to an exemplary embodiment. The application environment may include a terminal 100 and a server 200 .

[0030] In an optional embodiment, the terminal 100 can be used to provide multimedia resource generation services to any user. Specifically, the terminal 100 may include, but is not limited to, electronic devices such as smartphones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on these electronic devices, such as applications. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, Windows, etc.

[0031] In an optional embodiment, the server 200 can provide background services for the terminal 100. Specifically, the server 200 can be used to pre-train a multimedia resource generation model and provide multimedia resource generation services to the terminal 100 based on the trained multimedia resource generation model. Specifically, the server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0032] In addition, it should be noted that Figure 1 What is shown is only one application environment provided by the present disclosure. In actual application, other application environments may also be included.

[0033] In the embodiments of this specification, the terminal 100 and the server 200 may be directly or indirectly connected via wired or wireless communication, which is not limited in this disclosure.

[0034] Figure 2 This is a flowchart of a method for training a multimedia resource generation model according to an exemplary embodiment. The method can be applied to terminals, servers, etc. Figure 2 As shown, the method may include the following steps: In step S201, at least one preset multimedia resource corresponding to each of the plurality of currently generated description information is obtained; In a specific embodiment, the multiple currently generated description information may be multiple resource generation description information corresponding to the current training round during iterative training (multi-round training) of the multimedia resource generation model. Specifically, the resource generation description information may be description information of the multimedia resource to be generated. Specifically, the multiple currently generated description information may be randomly determined from a first preset generated description information set, or may be determined from the first preset generated description information set according to preset description information selection rules, for example, by sequentially selecting multiple currently generated description information that have not participated in training from the first preset generated description information set. Specifically, the first preset generated description information set may be a collection of resource generation description information used to train the multimedia resource generation model. Specifically, a multimedia resource is at least one resource capable of constituting multimedia. Exemplarily, the multimedia resource may include static resources such as images, or dynamic resources such as videos, that is, it may include at least one of images or videos. Optionally, the information format of the resource generation description information may include text, images, etc.

[0035] In a specific embodiment, the at least one preset multimedia resource corresponding to any currently generated description information (ie, any resource generated description information) may be at least one multimedia resource described by the currently generated description information. In a specific embodiment, the at least one preset multimedia resource may be one or more preset multimedia resources.

[0036] In step S203, based on the first preset noise information corresponding to the first current noisy time step, noise processing is performed on at least one preset multimedia resource to obtain at least one current noisy latent space feature corresponding to each currently generated description information.

[0037] In a specific embodiment, the first current noise addition time step can be the noise addition time step of the current training round in the iterative training process of the multimedia resource generation model. Specifically, the noise addition time step can represent the intensity of the noise addition; specifically, the first current noise addition time step is the noise addition time step corresponding to the current training round in the first preset time step interval; in each round of training, the current noise addition time step can be selected from the first preset time step interval, optionally, it can be selected randomly, or from small to large, or from large to small; specifically, the current noise addition time step corresponding to different training rounds can be different. Specifically, the first preset time step interval can be the noise addition time step interval in the iterative training process of the multimedia resource generation model. Specifically, the first preset noise information corresponding to the first current noise addition time step can be the noise information of the noise intensity corresponding to the first current noise addition time step. Specifically, the noise information can include but is not limited to Gaussian noise information, etc.

[0038] In a specific embodiment, the noise addition process is often performed in the latent space. Accordingly, the noise addition process for at least one preset multimedia resource based on the first preset noise information corresponding to the first current noise addition time step to obtain at least one current noisy latent space feature corresponding to each currently generated description information may include: performing noise addition process on the original latent space feature corresponding to at least one preset multimedia resource based on the first preset noise information corresponding to the first current noise addition time step to obtain at least one current noisy latent space feature corresponding to each currently generated description information. Specifically, the original latent space feature corresponding to any preset multimedia resource may be a resource feature (latent space vector) obtained by encoding the preset multimedia resource.

[0039] In step S205, at least one current noisy latent space feature, the first current noisy time step, and each current generated description information are input into a preset reward model for reward prediction to obtain first predicted reward data.

[0040] In a specific embodiment, the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource. Specifically, the first predicted reward data can be a numerical value used to measure the quality of the corresponding preset multimedia resource.

[0041] In a specific embodiment, the preset reward model may be a reward model that can directly process latent space features. Specifically, the preset reward model may be used to predict reward data (ie, predict the quality of multimedia resources).

[0042] In an optional embodiment, the preset reward model includes a latent space diffusion model and a reward prediction model; accordingly, the step of inputting at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the preset reward model for reward prediction to obtain the first predicted reward data may include: Inputting at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a latent space diffusion model to perform resource quality feature extraction processing to obtain a resource quality feature; The resource quality characteristics are input into the reward prediction model to perform reward prediction and obtain the first predicted reward data.

[0043] In a specific embodiment, latent diffusion models (LDMs) are an improvement on diffusion models (DMs). While diffusion models perform diffusion and denoising in the original image pixel space, LDMs perform diffusion and denoising in a low-dimensional latent space (i.e., they can directly process latent space features), reducing resource consumption and accelerating model training. Latent space features are typically encoded and decoded using a variational auto-encoder (VAE).

[0044] In a specific embodiment, the resource quality feature may be a quality feature corresponding to each preset multimedia resource.

[0045] In a specific embodiment, a reward prediction model can be used to convert resource quality features into corresponding reward data; specifically, the model structure of the reward prediction model can be set in combination with actual applications. Exemplarily, the reward prediction model can include an average pooling layer and a fully connected layer connected in sequence. Exemplarily, for example, a multimedia resource is preset to be a 1024*1024 image. After noise processing in the latent space, it becomes a latent space feature of size 64*64 (current noisy latent space feature). Further, assuming that the number of channels of the latent space diffusion model is C, the resource quality feature can be a 64*64*C dimensional feature (feature vector); further, after processing by the average pooling layer in the reward prediction model, a 1*C feature can be obtained; then, the 1*C feature is input into the fully connected layer in the reward prediction model to obtain the first predicted reward data.

[0046] In the above embodiment, the latent space diffusion model and the reward prediction model are used as the preset reward model. Each currently generated description information, at least one current noisy latent space feature corresponding to each currently generated description information, and the corresponding first current noisy time step can be directly input into the preset reward model. There is no need to convert the current noisy latent space feature into an image and then decode it when predicting the reward. This greatly reduces the consumption of computing resources, improves the system performance and model training efficiency, and directly processes the latent space features. This can avoid the situation where high-noise latent space features cannot restore the real image, resulting in the inability to effectively supervise high-noise areas, greatly improving the reward discrimination ability of the multimedia resource generation model to be trained under different noise conditions.

[0047] In an optional embodiment, the above method further includes the step of training a preset reward model, specifically, Figure 3 As shown, the following steps may be included: In step S301, at least one sample multimedia resource corresponding to each of a plurality of sample generation description information and preset quality data corresponding to the at least one sample multimedia resource are obtained; In step S303, based on the second preset noise information corresponding to the second current noise addition time step, at least one sample multimedia resource is subjected to noise addition processing to obtain at least one sample noisy latent space feature corresponding to each sample generation description information; In step S305, at least one sample noisy latent space feature, the second current noisy time step, and multiple sample generation description information are input into the reward model to be trained to perform reward prediction, thereby obtaining second predicted reward data; In step S307, quality prediction loss information is determined based on the preset quality data and the second prediction reward data; In step S309 , the reward model to be trained is iteratively trained based on the quality prediction loss information to obtain a preset reward model.

[0048] In a specific embodiment, a plurality of sample generation description information may be resource generation description information for reward model training in the current training round, and accordingly, the at least one sample multimedia resource corresponding to each sample generation description information may be at least one multimedia resource described by the sample generation description information. In a specific embodiment, the at least one sample multimedia resource may be one or more sample multimedia resources. Specifically, a plurality of sample generation description information may be randomly determined from the second preset generation description information set, or a plurality of sample generation description information may be determined from the second preset generation description information set according to a preset description information selection rule, for example, a plurality of sample generation description information that has not participated in training may be sequentially selected from the second preset generation description information set. Specifically, the second preset generation description information set may be a collection of resource generation description information for training a preset reward model.

[0049] In a specific embodiment, the preset quality data corresponding to the at least one sample multimedia resource corresponding to each sample generation description information may be data used to measure the quality of the at least one sample multimedia resource.

[0050] In an optional embodiment, when each sample generated description information corresponds to at least one sample multimedia resource, the preset quality data corresponding to the at least one sample multimedia resource may be data used to measure the quality of each of the at least one sample multimedia resources.

[0051] In a specific embodiment, preset quality data used to measure the quality of multimedia resources can be determined based on at least one resource measurement dimension. Specifically, the at least one resource measurement dimension can include at least one of the correlation between the multimedia resource and the corresponding resource-generated description information, the user's overall satisfaction with the multimedia resource, and the quality data of the resource's visual attributes. Optionally, the correlation can be pre-set by relevant personnel, or can be obtained by extracting features (feature vectors) from the multimedia resource and the corresponding resource-generated description information and using the distance between the features as the corresponding correlation. Specifically, the distance between the features can include, but is not limited to, Euclidean distance, Manhattan distance, etc. Specifically, the user's overall satisfaction with the multimedia resource can be pre-set by relevant personnel. Specifically, the resource visual attributes can be attributes reflecting the visual characteristics of the resource. For example, if the multimedia resource is an image, the resource visual attributes can include at least one of image clarity and image saturation. If the multimedia resource is a video, the resource visual attributes can include at least one of video frame clarity, video frame saturation, and smoothness between video frames. Optionally, the quality data of the resource visual attributes can be pre-set by relevant personnel or can be identified using artificial intelligence technology.

[0052] In a specific embodiment, when at least one resource measurement dimension comprises multiple resource measurement dimensions, a weighted sum of the quality data corresponding to each sample multimedia resource across the multiple resource measurement dimensions (e.g., the aforementioned correlation, overall satisfaction, and quality data of the resource's own attributes) can be performed to obtain the corresponding preset quality data. Specifically, the weight of each resource measurement dimension can be set based on actual application requirements.

[0053] In the above embodiment, when each sample generates description information corresponding to at least one sample multimedia resource, the data for measuring the quality of at least one sample multimedia resource is used as the corresponding preset quality data, which can intuitively and clearly reflect the quality of the resources, and can also facilitate subsequent models to intuitively learn the quality information of multimedia resources of different qualities.

[0054] In practical applications, in some scenarios where manual quality data labeling is required, direct manual quality data setting is often subjective and prone to errors in quality data setting. However, manual quality ranking of multiple images and other multimedia resources is often relatively accurate.

[0055] In an optional embodiment, when at least one sample multimedia resource corresponding to each sample generated description information is a plurality of sample multimedia resources, the preset quality data corresponding to the above-mentioned plurality of sample multimedia resources may be quality sorting information corresponding to the plurality of sample multimedia resources, and the quality sorting information is information that sorts the plurality of sample multimedia resources in ascending or descending order according to the quality of the resources corresponding to the plurality of sample multimedia resources.

[0056] In the above embodiment, the quality ranking information corresponding to multiple sample multimedia resources is used as the preset quality data corresponding to the multiple sample multimedia resources, which can effectively ensure the accuracy of the preset quality data in representing the quality of multiple sample multimedia resources. Furthermore, the quality ranking information of different sample multimedia resources can be combined to allow the reward model to more accurately learn the quality differences between different resources, thereby better improving the resource quality prediction accuracy of the reward model.

[0057] In a specific embodiment, the second current noise addition time step may be the noise addition time step of the current training round during the iterative training of the preset reward model. Specifically, the noise addition time step may represent the intensity of the noise addition. Specifically, in each round of training, the current noise addition time step may be selected from the second preset time step interval. Optionally, it may be selected randomly, or from small to large, or from large to small. Specifically, the current noise addition time step corresponding to different training rounds may be different. Specifically, the second preset time step interval may be the noise addition time step interval during the iterative training of the preset reward model. Specifically, the second preset noise information corresponding to the second current noise addition time step may be the noise information of the noise intensity corresponding to the second current noise addition time step. Specifically, the noise information may include but is not limited to Gaussian noise information, etc.

[0058] In a specific embodiment, based on the second preset noise information corresponding to the second current noise addition time step, at least one sample multimedia resource is noised to obtain a specific refinement of at least one sample noisy latent space feature corresponding to each sample generated description information. Please refer to the above-mentioned first preset noise information corresponding to the first current noise addition time step, at least one preset multimedia resource is noised to obtain a specific refinement of at least one current noisy latent space feature corresponding to each currently generated description information, which will not be repeated here.

[0059] In a specific embodiment, the second predicted reward data may be used to measure the quality of the corresponding sample multimedia resource; specifically, the second predicted reward data may be a numerical value used to measure the quality of the corresponding sample multimedia resource.

[0060] In an optional embodiment, when the at least one sample multimedia resource corresponding to each sample generation description information is a plurality of sample multimedia resources, and the preset quality data corresponding to the plurality of sample multimedia resources is quality ranking information corresponding to the plurality of sample multimedia resources, the determining of the quality prediction loss information based on the preset quality data and the second prediction reward data may include: Determine, from the plurality of sample multimedia resources corresponding to each sample generation description information, a positive sample multimedia resource corresponding to each sample generation description information and a negative sample multimedia resource corresponding to each sample generation description information according to the preset quality data; Quality prediction loss information is determined according to the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information.

[0061] In a specific embodiment, the resource quality corresponding to the positive sample multimedia resource is better than the resource quality corresponding to the negative sample multimedia resource. Optionally, when the multiple sample multimedia resources corresponding to each sample generation description information are two sample multimedia resources, and the preset quality data (quality sorting information) corresponding to the two sample multimedia resources is information that is sorted in ascending order according to the quality quality of the two sample multimedia resources (i.e., sorted from inferior to superior), the sample multimedia resource with the first corresponding quality sorting information can be used as the negative sample multimedia resource, and the sample multimedia resource with the second corresponding quality sorting information can be used as the positive sample multimedia resource. Conversely, if the preset quality data (quality sorting information) corresponding to the two sample multimedia resources is information that is sorted in descending order according to the quality quality of the two sample multimedia resources (i.e., sorted from superior to inferior), the sample multimedia resource with the second corresponding quality sorting information can be used as the negative sample multimedia resource, and the sample multimedia resource with the first corresponding quality sorting information can be used as the positive sample multimedia resource.

[0062] In an optional embodiment, the multiple sample multimedia resources corresponding to the description information generated for each sample are at least three sample multimedia resources. Two sample multimedia resources can be randomly selected from the at least three sample multimedia resources to determine the positive and negative sample multimedia resources, or the preset resource selection rules can be followed, for example, the preset quality data corresponding to the two sample multimedia resources ranked first and last are selected to determine the positive and negative sample multimedia resources; optionally, when the preset quality data (quality sorting information) corresponding to the at least three sample multimedia resources is information that is sorted in ascending order according to the quality of the resources corresponding to the two sample multimedia resources, the sample multimedia resource with the lower ranking among the two selected sample multimedia resources is used as the positive sample multimedia resource, and the sample multimedia resource with the higher ranking is used as the negative sample multimedia resource; when the preset quality data (quality sorting information) corresponding to the at least three sample multimedia resources is information that is sorted in descending order according to the quality of the resources corresponding to the two sample multimedia resources, the sample multimedia resource with the higher ranking among the two selected sample multimedia resources can be used as the positive sample multimedia resource, and the sample multimedia resource with the lower ranking is used as the negative sample multimedia resource.

[0063] In a specific embodiment, the first preset loss function can be combined in the process of determining the quality prediction loss information based on the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information; specifically, the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information can be input into the first preset loss function to obtain the quality prediction loss information; specifically, the quality prediction loss information can characterize the prediction accuracy of the resource quality of the reward model to be trained.

[0064] In a specific embodiment, the above-mentioned first preset loss function can be set in combination with actual applications; specifically, the first preset loss function is a ranking loss function, for example, a binary cross entropy loss function.

[0065] In the above embodiment, when at least one sample multimedia resource corresponding to each sample generation description information is a plurality of sample multimedia resources, and the preset quality data corresponding to the plurality of sample multimedia resources is the quality ranking information corresponding to the plurality of sample multimedia resources, the preset quality data (quality ranking information) is first combined to determine the positive sample multimedia resource corresponding to each sample generation description information and the negative sample multimedia resource corresponding to each sample generation description information from the plurality of sample multimedia resources corresponding to each sample generation description information, and the quality prediction loss information is determined based on the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information. Subsequently, based on the quality prediction loss information, with the goal of making the positive and negative samples far enough apart, the reward model can more accurately learn the quality differences between different resources, thereby improving the resource quality prediction accuracy of the reward model.

[0066] In an optional embodiment, if each sample generation description information corresponds to at least one sample multimedia resource, and the preset quality data corresponding to the at least one sample multimedia resource corresponding to each sample generation description information is data used to measure the quality of the at least one sample multimedia resource, the quality prediction loss information can be determined in combination with the second preset loss function; specifically, the preset quality data corresponding to each sample multimedia resource and the corresponding second predicted reward data can be input into the second preset loss function to obtain the quality prediction loss information. Specifically, the second preset loss function can be set in combination with actual applications, such as a cross-entropy loss function, etc., and subsequently based on the quality prediction loss information, in order to make the preset quality data and the corresponding second predicted reward data close enough, the reward model can learn the resource quality, thereby improving the resource quality prediction accuracy of the reward model.

[0067] In a specific embodiment, based on the quality prediction loss information, iteratively training the reward model to be trained to obtain the preset reward model may include: updating the model parameters of the reward model to be trained according to the quality prediction loss information, and performing the next training round based on the reward model to be trained after the updated model parameters, that is, repeating the above steps S301-step S307, and the step of updating the model parameters of the reward model to be trained according to the quality prediction loss information, until the first preset convergence condition is met, and using the reward model to be trained when the first preset convergence condition is met as the preset reward model.

[0068] In a specific embodiment, the first preset convergence condition can be set in combination with actual applications, for example, the quality prediction loss information is less than the first preset threshold, or the difference between the pairwise quality prediction loss information obtained in the most recent first preset round is less than the second preset threshold, or the training round reaches the third preset threshold, etc.

[0069] In the above embodiment, during the iterative training of the reward model, the sample multimedia resource is first noised in combination with the second preset noise information corresponding to the second current noise adding time step to obtain at least one sample noised latent space feature corresponding to each sample generation description information. Then, the at least one sample noised latent space feature, the second current noise adding time step and multiple sample generation description information are input into the reward model to be trained for reward prediction to obtain second predicted reward data. Based on the preset quality data and the second predicted reward data, the quality prediction loss information is determined. Based on the quality prediction loss information, the reward model to be trained is iteratively trained to obtain the preset reward model. This allows the reward model to learn resource quality and effectively ensure the resource quality prediction accuracy of the reward model. Furthermore, the parameters of the multimedia resource generation model can be optimized during the iterative training of the multimedia resource generation model to maximize the reward score of the generated resources and ensure the multimedia resource generation effect of the trained multimedia resource generation model.

[0070] In step S207, reward loss information is determined based on the first predicted reward data.

[0071] In a specific embodiment, in the process of determining the reward loss information based on the first predicted reward data, a third preset loss function can be combined; specifically, the third preset loss function can be set in combination with actual applications; for example, a Sigmoid loss function, etc.

[0072] In a specific embodiment, the first predicted reward data can be input into a third preset loss function to obtain reward loss information; specifically, the reward loss information can represent the denoising performance of the current multimedia resource generation model to be trained in the latent space.

[0073] In step S209, denoising reasoning is performed on the multimedia resource generation model to be trained based on at least one current noisy latent space feature, each current generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; In a specific embodiment, the multimedia resource generation model to be trained may be a generative model to be trained. The specific model structure may be set in combination with actual applications. For example, the multimedia resource generation model to be trained may be a diffusion model (DM) to be trained.

[0074] In an optional embodiment, performing denoising inference on the multimedia resource generation model to be trained based on at least one current noisy latent space feature, each currently generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained may include: Inputting at least one current noisy latent space feature and each currently generated description information into the multimedia resource generation model to be trained for denoising, thereby obtaining first noise prediction information; Inputting at least one current noisy latent space feature and multiple current generated description information into a reference model corresponding to the multimedia resource generation model to be trained for denoising, obtaining second noise prediction information; Noise prediction loss information is determined according to the first noise prediction information, the second noise prediction information, and the first preset noise information.

[0075] In a specific embodiment, the reference model corresponding to the multimedia resource generation model to be trained can be the multimedia resource generation model to be trained in the initial state (that is, the multimedia resource generation model to be trained that has not been iteratively updated), and the reference model remains unchanged during the iterative training process of the multimedia resource generation model to be trained.

[0076] In a specific embodiment, the process of generating multimedia resources by the multimedia resource generation model to be trained (i.e., the process of performing denoising processing) may include a process of predicting noise information and a process of removing noise information based on the predicted noise information; during the training phase of the multimedia resource generation model to be trained, only the predicted noise information (i.e., noise prediction information) may be output.

[0077] In a specific embodiment, the first noise prediction information may be noise information predicted by the multimedia resource generation model to be trained; the second noise prediction information may be noise information predicted by a reference model corresponding to the multimedia resource generation model to be trained.

[0078] In a specific embodiment, a fourth preset loss function may be incorporated into the process of determining noise prediction loss information based on the first noise prediction information, the second noise prediction information, and the first preset noise information. Specifically, the fourth preset loss function may be configured based on actual applications, such as, for example, a mean square error loss function. Specifically, the first noise prediction information, the second noise prediction information, and the first preset noise information may be input into the fourth preset loss function to obtain noise prediction loss information. Specifically, the noise prediction loss information may represent the noise prediction performance (i.e., denoising performance) of the current multimedia resource to be trained.

[0079] In the above embodiment, at least one current noisy latent space feature and each currently generated description information are input into the multimedia resource generation model to be trained for denoising processing to obtain first noise prediction information, and at least one current noisy latent space feature and multiple currently generated description information are input into the reference model corresponding to the multimedia resource generation model to be trained for denoising processing to obtain second noise prediction information, and based on the first noise prediction information, the second noise prediction information and the first preset noise information, the noise prediction loss information is determined. The noise information predicted by the reference model can be combined to guide the multimedia resource generation model to be trained to retain the original performance during denoising, and the model can also learn diverse features.

[0080] In step S211 , the multimedia resource generation model to be trained is iteratively trained based on the noise prediction loss information and the reward loss information to obtain a multimedia resource generation model.

[0081] In a specific embodiment, the iterative training of the multimedia resource generation model to be trained based on the noise prediction loss information and the reward loss information to obtain the multimedia resource generation model may include: Perform weighted summation of noise prediction loss information and reward loss information to obtain target loss information; Based on the target loss information, the multimedia resource generation model to be trained is iteratively trained to obtain a multimedia resource generation model.

[0082] In a specific embodiment, the weights of the noise prediction loss information and the reward loss information can be set in combination with actual applications.

[0083] In a specific embodiment, based on the target loss information, the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model, which may include updating the model parameters of the multimedia resource generation model to be trained according to the target loss information, and performing the next training round based on the multimedia resource generation model to be trained after the model parameters are updated, that is, repeating the above steps S201-S209, and performing weighted summation of the noise prediction loss information and the reward loss information to obtain the target loss information, and updating the model parameters of the multimedia resource generation model to be trained according to the target loss information, until the second preset convergence condition is met, and the multimedia resource generation model to be trained that meets the second preset convergence condition is used as the multimedia resource generation model.

[0084] In a specific embodiment, the second preset convergence condition can be set in combination with the actual application, for example, the target loss information is less than the fourth preset threshold, or the difference between the two target loss information obtained in the most recent second preset round is less than the fifth preset threshold, or the training round reaches the sixth preset threshold, etc.

[0085] It can be seen from the technical solutions provided by the above embodiments of the present application that, in the iterative training process of the multimedia resource generation model, the present application performs noise processing in the latent space on at least one preset multimedia resource corresponding to each of the multiple currently generated description information based on the first preset noise information corresponding to the first current noise adding time step, and obtains at least one current noisy latent space feature corresponding to each currently generated description information, and directly inputs the first current noisy time step, each currently generated description information and at least one current noisy latent space feature into the preset reward model for reward prediction, and can perform reward prediction without restoring multimedia resources such as images, thereby simplifying the model training steps and improving the model training efficiency (compared to the existing technology based on reinforcement learning, the training efficiency can be increased by 3 times), and in the iterative training process of the model, different noise adding time steps can be selected from the first preset time step interval, It effectively avoids the situation where high-noise latent space features cannot restore the real image and cannot effectively supervise the high-noise area, which can greatly improve the reward discrimination ability of the multimedia resource generation model to be trained under different noise conditions, and determines the reward loss information by combining the first predicted reward data obtained by reward prediction for measuring the quality of the corresponding preset multimedia resource; and based on the reward loss information and at least one current noisy latent space feature, each current generated description information and the first preset noise information, denoising reasoning is performed on the multimedia resource generation model to be trained, and the noise prediction loss information corresponding to the multimedia resource generation model to be trained is obtained, and the multimedia resource generation model to be trained is iteratively trained to obtain the multimedia resource generation model, which can greatly improve the model training effect, thereby improving the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.

[0086] Based on the training method of the multimedia resource generation model provided in the embodiment of the present application, a multimedia resource generation method provided in the embodiment of the present application is introduced. Figure 4 is a flowchart of a method for generating multimedia resources according to an exemplary embodiment. The method can be applied to a terminal, a server, etc. Figure 4 As shown, the method may include the following steps: In step S401, target generation description information and preset noisy multimedia resources are obtained.

[0087] In a specific embodiment, the target generation description information may be description information of a multimedia resource to be generated (resource generation description information); the target generation description information may include text or an image. In a specific embodiment, the preset noisy multimedia resource may be a multimedia resource generated based on preset noise; for example, a multimedia resource such as a Gaussian noise image generated based on Gaussian noise.

[0088] In step S403, the target generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model to perform multimedia resource generation processing, and obtain the target multimedia resource corresponding to the target generation description information.

[0089] In a specific embodiment, the target generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model for multimedia resource generation processing, and obtaining the target multimedia resource corresponding to the target generation description information may include: inputting the target generation description information and the preset noisy multimedia resource into the multimedia resource generation model, the multimedia resource generation model performs noise prediction based on the target generation description information and the preset noisy multimedia resource to obtain target noise prediction information, and denoising the preset noisy multimedia resource based on the target noise prediction information to obtain the target multimedia resource.

[0090] In addition, it should be noted that the multiple in the embodiments of the present application can be at least two.

[0091] It can be seen from the technical solutions provided in the above embodiments of this specification that, in the multimedia resource generation process, the target generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model for multimedia resource generation processing. During the iterative training of the multimedia resource generation model, based on the first preset noise information corresponding to the first current noisy time step, at least one preset multimedia resource corresponding to each of the multiple current generation description information is noised in the latent space to obtain at least one current noisy latent space feature corresponding to each current generation description information, and the first current noisy time step, each current generation description information and at least one current noisy latent space feature are directly input into the preset reward model for reward prediction. Reward prediction can be performed without restoring multimedia resources such as images, thereby simplifying the model training steps and improving the model training efficiency. In the iterative training of the model, Selecting different noise adding time steps in the step interval effectively avoids the situation where high-noise latent space features cannot restore the real image and cannot effectively supervise the high-noise area, which can greatly improve the reward discrimination ability of the multimedia resource generation model to be trained under different noise conditions, and combines the first predicted reward data obtained by reward prediction to measure the quality of the corresponding preset multimedia resource to determine the reward loss information; and based on the reward loss information and at least one current noisy latent space feature, each current generated description information and the first preset noise information, the multimedia resource generation model to be trained is subjected to denoising reasoning, and the obtained noise prediction loss information corresponding to the multimedia resource generation model to be trained is used together to iteratively train the multimedia resource generation model to obtain a multimedia resource generation model, which can greatly improve the model training effect, thereby effectively improving the denoising performance of the multimedia resource generation model and the generation effect of multimedia resources.

[0092] Figure 5 FIG1 is a block diagram of a training device for a multimedia resource generation model according to an exemplary embodiment. Figure 5 , the device comprises: The multimedia resource acquisition module 510 is configured to acquire at least one preset multimedia resource corresponding to each of the plurality of currently generated description information; A first noise processing module 520 is configured to perform noise processing on at least one preset multimedia resource based on first preset noise information corresponding to a first current noise adding time step, to obtain at least one current noisy latent space feature corresponding to each currently generated description information; A first reward prediction module 530 is configured to perform reward prediction by inputting at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a preset reward model to obtain first predicted reward data, where the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource; The reward loss information determination module 540 is configured to determine reward loss information based on the first predicted reward data; The denoising inference module 550 is configured to perform denoising inference on the multimedia resource generation model to be trained based on at least one current noisy latent space feature, each current generated description information, and the first preset noise information, to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; The first model iterative training module 560 is configured to perform iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information and the reward loss information to obtain the multimedia resource generation model.

[0093] In an optional embodiment, the preset reward model includes a latent space diffusion model and a reward prediction model; the first reward prediction module 530 includes: a resource quality feature extraction unit configured to perform resource quality feature extraction processing by inputting at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a latent space diffusion model to obtain a resource quality feature; The reward prediction unit is configured to input the resource quality characteristics into the reward prediction model to perform reward prediction and obtain first predicted reward data.

[0094] In an optional embodiment, the above device further includes: A training data acquisition module is configured to acquire at least one sample multimedia resource corresponding to each of a plurality of sample generation description information and preset quality data corresponding to the at least one sample multimedia resource; a noise processing module configured to perform noise processing on at least one sample multimedia resource based on second preset noise information corresponding to a second current noise adding time step, and obtain at least one sample noisy latent space feature corresponding to each sample generated description information; A second reward prediction module is configured to perform reward prediction by inputting at least one sample noisy latent space feature, a second current noisy time step, and multiple sample generation description information into the to-be-trained reward model to obtain second predicted reward data, where the second predicted reward data is used to measure the quality of the corresponding sample multimedia resource; a quality prediction loss information determining module configured to determine quality prediction loss information based on preset quality data and second prediction reward data; The second model iterative training module is configured to perform iterative training on the reward model to be trained based on the quality prediction loss information to obtain a preset reward model.

[0095] In an optional embodiment, when at least one sample multimedia resource corresponding to each sample generated description information is a plurality of sample multimedia resources, the preset quality data corresponding to the plurality of sample multimedia resources is quality sorting information corresponding to the plurality of sample multimedia resources, and the quality sorting information is information that sorts the plurality of sample multimedia resources in ascending or descending order according to the quality of the resources corresponding to the plurality of sample multimedia resources.

[0096] In an optional embodiment, the quality prediction loss information determination module includes: The positive and negative sample data acquisition unit is configured to determine, from the plurality of sample multimedia resources corresponding to each sample generation description information, a positive sample multimedia resource corresponding to each sample generation description information and a negative sample multimedia resource corresponding to each sample generation description information according to preset quality data; The quality prediction loss information determination unit is configured to determine the quality prediction loss information based on the second prediction reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second prediction reward data of the negative sample multimedia resource corresponding to each sample generation description information.

[0097] In an optional embodiment, when each sample generated description information corresponds to at least one sample multimedia resource, the preset quality data corresponding to the at least one sample multimedia resource is data used to measure the quality of each of the at least one sample multimedia resources.

[0098] In an optional embodiment, the denoising inference module 550 includes: A first denoising processing unit is configured to perform denoising processing on at least one current noisy latent space feature and each currently generated description information input into a multimedia resource generation model to be trained, thereby obtaining first noise prediction information; The second denoising processing unit is configured to perform denoising processing on the at least one currently noisy latent space feature and the multiple currently generated description information input into a reference model corresponding to the multimedia resource generation model to be trained, to obtain second noise prediction information; The noise prediction loss information determining unit is configured to determine the noise prediction loss information according to the first noise prediction information, the second noise prediction information and the first preset noise information.

[0099] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0100] Figure 6 FIG. 1 is a block diagram of a multimedia resource generation device according to an exemplary embodiment. Figure 6 , the device comprises: The data acquisition module 610 is configured to execute acquisition target generation description information and preset noise-added multimedia resources; The multimedia resource generation module 620 is configured to execute multimedia resource generation processing by inputting the target generation description information and the preset noisy multimedia resource into the multimedia resource generation model obtained by the training method based on any of the above-mentioned multimedia resource generation models to obtain the target multimedia resource corresponding to the target generation description information.

[0101] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0102] Figure 7 This is a block diagram of an electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a training method for a multimedia resource generation model or a multimedia resource generation method is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc. Figure 8 This is a block diagram of another electronic device for training a multimedia resource generation model or generating multimedia resources according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The electronic device includes a processor, a memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a training method for a multimedia resource generation model or a multimedia resource generation method is implemented. Those skilled in the art will understand that Figure 7 or Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present disclosure, and does not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, an electronic device is also provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a training method for a multimedia resource generation model or a multimedia resource generation method as in an embodiment of the present disclosure.

[0103] In an exemplary embodiment, a computer-readable storage medium is also provided. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the training method of the multimedia resource generation model or the multimedia resource generation method in the embodiment of the present disclosure. In an exemplary embodiment, a computer program product including instructions is further provided. When the computer program product is run on a computer, the computer is caused to execute the multimedia resource generation model training method or multimedia resource generation method in the embodiments of the present disclosure.

[0104] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, which can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0105] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0106] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A training method for a multimedia resource generation model, characterized in that: include: Obtain at least one preset multimedia resource corresponding to each of the multiple currently generated description information; Noising the at least one preset multimedia resource based on first preset noise information corresponding to a first current noisy time step, to obtain at least one current noisy latent space feature corresponding to each currently generated description information, wherein the first current noisy time step is a noisy time step corresponding to a current training round in the first preset time step interval; Inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a preset reward model to perform reward prediction, obtaining first predicted reward data, where the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource; Determining reward loss information based on the first predicted reward data; Performing denoising reasoning on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each currently generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; Based on the noise prediction loss information and the reward loss information, the multimedia resource generation model to be trained is iteratively trained to obtain a multimedia resource generation model.

2. The method for training a multimedia resource generation model according to claim 1, wherein: The preset reward model includes a latent space diffusion model and a reward prediction model; the step of inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the preset reward model for reward prediction to obtain first predicted reward data includes: Inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into the latent space diffusion model to perform resource quality feature extraction processing to obtain a resource quality feature; The resource quality characteristics are input into the reward prediction model to perform reward prediction to obtain the first predicted reward data.

3. The training method for a multimedia resource generation model according to claim 1, characterized in that: The method further comprises: Acquire at least one sample multimedia resource corresponding to each of the plurality of sample generation description information and preset quality data corresponding to the at least one sample multimedia resource; performing noise processing on the at least one sample multimedia resource based on second preset noise information corresponding to the second current noise addition time step, to obtain at least one sample noisy latent space feature corresponding to each sample generation description information; Inputting the at least one sample noisy latent space feature, the second current noisy time step, and the multiple sample generation description information into the to-be-trained reward model to perform reward prediction, obtaining second predicted reward data, where the second predicted reward data is used to measure the quality of the corresponding sample multimedia resource; determining quality prediction loss information based on the preset quality data and the second prediction reward data; Based on the quality prediction loss information, the reward model to be trained is iteratively trained to obtain the preset reward model.

4. The method for training a multimedia resource generation model according to claim 3, wherein: In the case where at least one sample multimedia resource corresponding to each sample generated description information is a plurality of sample multimedia resources, the preset quality data corresponding to the plurality of sample multimedia resources is the quality sorting information corresponding to the plurality of sample multimedia resources, and the quality sorting information is information that sorts the plurality of sample multimedia resources in ascending or descending order according to the quality of the resources corresponding to the plurality of sample multimedia resources.

5. The method for training a multimedia resource generation model according to claim 4, wherein: The determining of quality prediction loss information based on the preset quality data and the second prediction reward data includes: Determining, according to the preset quality data, a positive sample multimedia resource corresponding to each sample generation description information and a negative sample multimedia resource corresponding to each sample generation description information from the multiple sample multimedia resources corresponding to each sample generation description information; The quality prediction loss information is determined according to the second predicted reward data of the positive sample multimedia resource corresponding to each sample generation description information and the second predicted reward data of the negative sample multimedia resource corresponding to each sample generation description information.

6. The method for training a multimedia resource generation model according to claim 3, wherein: In the case where each sample generated description information corresponds to at least one sample multimedia resource, the preset quality data corresponding to the at least one sample multimedia resource is data used to measure the quality of each resource of the at least one sample multimedia resource.

7. The method for training a multimedia resource generation model according to any one of claims 1 to 6, characterized in that: The performing denoising reasoning on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each current generated description information, and the first preset noise information to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained includes: Inputting the at least one current noisy latent space feature and each currently generated description information into the multimedia resource generation model to be trained for denoising, to obtain first noise prediction information; Inputting the at least one current noisy latent space feature and the multiple current generated description information into a reference model corresponding to the multimedia resource generation model to be trained for denoising, to obtain second noise prediction information; The noise prediction loss information is determined according to the first noise prediction information, the second noise prediction information, and the first preset noise information.

8. A method for generating multimedia resources, characterized in that: include: Obtain target generation description information and preset noise-added multimedia resources; The target generation description information and the preset noisy multimedia resource are input into the multimedia resource generation model obtained based on the training method of the multimedia resource generation model according to any one of claims 1 to 7 to perform multimedia resource generation processing to obtain the target multimedia resource corresponding to the target generation description information.

9. A training device for a multimedia resource generation model, characterized in that: include: A multimedia resource acquisition module is configured to acquire at least one preset multimedia resource corresponding to each of the plurality of currently generated description information; a first noise processing module configured to perform noise processing on the at least one preset multimedia resource based on first preset noise information corresponding to a first current noise adding time step, to obtain at least one current noisy latent space feature corresponding to each currently generated description information; A first reward prediction module is configured to perform reward prediction by inputting the at least one current noisy latent space feature, the first current noisy time step, and each current generated description information into a preset reward model to obtain first predicted reward data, where the first predicted reward data is used to measure the quality of the corresponding preset multimedia resource; a reward loss information determination module, configured to determine reward loss information based on the first predicted reward data; a denoising inference module configured to perform denoising inference on the multimedia resource generation model to be trained based on the at least one current noisy latent space feature, each currently generated description information, and the first preset noise information, to obtain noise prediction loss information corresponding to the multimedia resource generation model to be trained; The first model iterative training module is configured to perform iterative training on the multimedia resource generation model to be trained based on the noise prediction loss information and the reward loss information to obtain a multimedia resource generation model.

10. A multimedia resource generating device, characterized in that: include: A data acquisition module is configured to execute acquisition target generation description information and preset noise-added multimedia resources; The multimedia resource generation module is configured to execute multimedia resource generation processing by inputting the target generation description information and the preset noisy multimedia resource into a multimedia resource generation model obtained by the training method of the multimedia resource generation model according to any one of claims 1 to 7, so as to obtain the target multimedia resource corresponding to the target generation description information.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of the multimedia resource generation model according to any one of claims 1 to 7 or the multimedia resource generation method according to claim 8.

12. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method for a multimedia resource generation model as described in any one of claims 1 to 7 or the multimedia resource generation method as described in claim 8.

13. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method for training a multimedia resource generation model according to any one of claims 1 to 7 or the method for generating multimedia resources according to claim 8.

Citation Information

Patent Citations

  • Model training method, audio generation method, computer equipment and storage medium

    CN118098268A

  • Method and system for supervising pixel-by-pixel loss based on diffusion model

    CN118351545A

  • Image generation model training method and device, electronic equipment and storage medium

    CN119919755A

  • Training method of image generation model, image generation method, device, equipment, storage medium and program product

    CN120219201A

  • Method and apparatus for determining picture generation model, method and apparatus for picture generation, computing device, storage medium, and program product

    WO2024239755A1