Model training method, data generation method and related equipment
By combining low-quality labeled and high-quality training data without labels, the diffusion model is trained, and noise data is generated using the noise-added module, the problem of high cost of training data of the diffusion model is solved and the generation of high-quality data is achieved.
Patent Information
- Application Number
- CN202410063854.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The high-quality labeled training data required for training diffusion models is expensive to obtain, and the labeling process consumes a lot of human resources and time.
By combining labeled low-quality training data and labelless high-quality training data, the diffusion model is trained, and noise data is generated using the noise addition module to learn the label characteristics and high-quality features of the data, and the denoising module is trained.
The training data acquisition cost of the diffusion model is reduced, while ensuring the quality of the generated data, realizing the generation of high-quality data.
Smart Images

Figure CN120338029A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular, to a model training method, a data generation method, and related devices. Background Art
[0002] Diffusion models are a class of generative models that learn the probability distribution of data by applying a series of random perturbations to the data. These perturbations are usually Gaussian noise that gradually destroys the data until it becomes pure noise. The generation process is then reversed by applying a series of denoising steps that gradually recover the original data from the pure noise.
[0003] Diffusion models can sample from the learned distribution to generate realistic and diverse data samples such as images, text, audio, and video. Diffusion models can also generate data related to a given condition based on additional information such as text or audio. For example, generating an image that matches the description of a text prompt; generating a video that matches the description of a text prompt; generating a 3D model that corresponds to the shape and appearance given in a text prompt description; generating an action sequence that matches the action or style given in a text prompt description; generating an image related to a given sound; generating a video synchronized with a given sound; generating dance movements harmonious with a given piece of music.
[0004] Training diffusion models requires relying on high-quality labeled training data. However, the cost of obtaining training data with high quality and rich annotations is very high. For example, taking motion data as an example of training data, an expensive motion capture system and professional operators are needed to collect high-quality motion data, and a large amount of human resources and time are also consumed for annotating the motion data. Summary of the Invention
[0005] The present application provides a model training method, a data generation method, and related devices, which can effectively reduce the acquisition cost of training data for diffusion models.
[0006] In a first aspect, a model training method is provided. This method can be executed by a model training device or by a chip in the model training device.
[0007] The above model training method includes the following steps: obtaining a plurality of first noise data corresponding to a plurality of first training data. The above first training data is labeled. The first noise data corresponding to the first training data is obtained by adding noise to the first training data for a first number of noise addition times. The first number of noise addition times is greater than the second number of noise addition times and less than or equal to the maximum number of noise addition times; the second number of noise addition times is related to the noise level of the plurality of first training data. Obtaining a plurality of second noise data corresponding to a plurality of second training data. The above second training data is unlabeled, and the noise in the first training data is higher than the noise in the second training data. The second noise data corresponding to the second training data is obtained by adding noise to the second training data for a third number of noise addition times; the third number of noise addition times is less than or equal to the second number of noise addition times and greater than or equal to one. Training the first denoising module of the diffusion model based on the plurality of first noise data and the plurality of second noise data to obtain a second denoising module.
[0008] The specific form of the above label can be text and / or audio. The label of the above first training data can be understood as the specified condition or given condition when generating relevant data based on the first training data. The above text can be a sentence or a word.
[0009] The above number of noise addition times can be understood as the noise addition time step or the degree of noise addition. For example, the larger the value of the first number of noise addition times, the longer the noise addition time for the first training data. On the contrary, the smaller the value of the first number of noise addition times, the less the noise addition time for the first training data. Another example is that the larger the value of the first number of noise addition times, the stronger the noise perturbation applied to the first training data. On the contrary, the smaller the value of the first number of noise addition times, the weaker the noise perturbation applied to the first training data.
[0010] The above second number of noise addition times is the number of noise addition times determined based on the noise level of the plurality of first training data.
[0011] It can be seen that in this solution, the first denoising module of the diffusion model is trained based on the plurality of first noise data and the plurality of second noise data to obtain a second denoising module. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, this solution can effectively reduce the acquisition cost of the training data of the diffusion model.
[0012] In a possible implementation of the first aspect, training the first denoising module of the diffusion model based on multiple first noise data and multiple second noise data to obtain a second denoising module specifically includes the following steps: Training the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data to obtain the second denoising module.
[0013] In this solution, the second denoising module can simultaneously learn the label features of the first training data and the high-quality data features of the second training data, enabling the second denoising module to achieve high-quality data generation with low training costs.
[0014] In a possible implementation of the first aspect, training the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data to obtain the second denoising module specifically includes the following steps: Predicting a first noise value of the first noise data through the first denoising module based on the label of the first training data, the first noise data, and the first noise addition time corresponding to the first noise data. The first noise value is a predicted value of the noise addition amount for the Nth noise addition, where N is the first noise addition time. Predicting a second noise value of the second noise data through the first denoising module based on the second noise data and the third noise addition time corresponding to the second noise data. The second noise value is a predicted value of the noise addition amount for the Mth noise addition, where M is the third noise addition time. Adjusting the first denoising module based on a first loss value and a second loss value. The first loss value is obtained based on the first noise value and a third noise value, and the third noise value is the true value of the noise addition amount for the Nth noise addition. The second loss value is obtained based on the second noise value and a fourth noise value, and the fourth noise value is the true value of the noise addition amount for the Mth noise addition.
[0015] In this solution, when training the first denoising module, by predicting the first noise value and the second noise value to obtain the first loss value and the second loss value, the first denoising module can then be adjusted based on the first loss value and the second loss value. Repeat the execution until the training stop condition is met, and then the above-mentioned second denoising module can be obtained.
[0016] In a possible implementation of the first aspect, adjusting the first denoising module based on the first loss value and the second loss value specifically includes the following steps: Determining a third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value. Adjusting the first denoising module based on the third loss value.
[0017] In this solution, the first weight value and the second weight value can be set to adjust the focus of model learning to meet different model training requirements.
[0018] In a possible implementation manner of the first aspect, the above model training method further includes the following steps: obtaining a plurality of first training data and a plurality of second training data. Performing noise addition processing on the plurality of first training data through the noise addition module of the diffusion model to obtain a plurality of first noise data. Performing noise addition processing on the plurality of second training data through the noise addition module to obtain a plurality of second noise data.
[0019] In this solution, the noise addition module of the diffusion model is used to obtain the first noise data and the second noise data.
[0020] In the second aspect, the present application further provides a data generation method, which can be executed by a data generation device or by a chip in the data generation device.
[0021] The above data generation method includes the following steps: obtaining indication information and random noise; the indication information is related to the first data. Generating the first data through the denoising module of the diffusion model based on the indication information and the random noise. The above denoising module is trained by using the model training method of the first aspect.
[0022] The above indication information is used to indicate the generation object or generation target of the denoising module, that is, the given condition.
[0023] In this solution, the denoising module can generate the first data based on the indication information and the random noise. The above denoising module is trained based on a plurality of first noise data and a plurality of second noise data. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, therefore, in this solution, the acquisition cost of the training data of the denoising module is low.
[0024] In the third aspect, the present application further provides a model training device, including an acquisition module and a training module.
[0025] The acquisition module is used to obtain a plurality of first noise data corresponding to a plurality of first training data. The above first training data is labeled, and the first noise data corresponding to the first training data is obtained after adding noise to the first training data for the first number of noise addition times. The first number of noise addition times is greater than the second number of noise addition times, and the first number of noise addition times is less than or equal to the maximum number of noise addition times; the second number of noise addition times is related to the noise level of the plurality of first training data.
[0026] The acquisition module is also used to acquire multiple second noise data corresponding to multiple second training data. The above-mentioned second training data has no labels, and the noise in the first training data is higher than the noise in the second training data. The second noise data corresponding to the second training data is obtained by adding noise to the second training data for the third noise addition times. The third noise addition times is less than or equal to the second noise addition times, and the third noise addition times is greater than or equal to one.
[0027] The training module is used to train the first denoising module of the diffusion model based on multiple first noise data and multiple second noise data to obtain a second denoising module.
[0028] It can be seen that in this solution, the model training device trains the first denoising module of the diffusion model based on multiple first noise data and multiple second noise data to obtain a second denoising module. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, therefore, the acquisition cost of the training data of the model training device in this solution is low.
[0029] In a possible implementation manner of the third aspect, the above-mentioned training module is specifically used to: train the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data, to obtain a second denoising module.
[0030] In a possible implementation manner of the third aspect, when the above-mentioned training module trains the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data, to obtain a second denoising module, it is specifically used to:
[0031] Based on the label of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data, predict the first noise value of the first noise data through the first denoising module. The above-mentioned first noise value is the predicted value of the noise addition amount for the Nth noise addition, and N is the first noise addition times. Based on the second noise data and the third noise addition times corresponding to the second noise data, predict the second noise value of the second noise data through the first denoising module. The above-mentioned second noise value is the predicted value of the noise addition amount for the Mth noise addition, and M is the third noise addition times. Adjust the first denoising module based on the first loss value and the second loss value. The first loss value is obtained based on the first noise value and the third noise value, and the third noise value is the true value of the noise addition amount for the Nth noise addition. The second loss value is obtained based on the second noise value and the fourth noise value, and the fourth noise value is the true value of the noise addition amount for the Mth noise addition.
[0032] In a possible implementation of the third aspect, when the above training module adjusts the first denoising module based on the first loss value and the second loss value, it is specifically configured to: determine a third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value. Adjust the first denoising module based on the third loss value.
[0033] In a possible implementation of the third aspect, the above acquisition module is further configured to acquire a plurality of first training data and a plurality of second training data. The above model training device further includes a processing module. The processing module is configured to perform noise addition processing on the plurality of first training data through the noise addition module of the diffusion model to obtain a plurality of first noise data. The processing module is further configured to perform noise addition processing on the plurality of second training data through the noise addition module to obtain a plurality of second noise data.
[0034] Fourth aspect, the present application further provides a data generation device, including an acquisition module and a generation module.
[0035] The acquisition module is configured to acquire indication information and random noise. The indication information is related to the first data.
[0036] The generation module is configured to generate the first data based on the indication information and the random noise through the denoising module of the diffusion model. The above denoising module is trained by using the model training method described in the first aspect.
[0037] In this solution, the data generation device can generate the first data based on the indication information and the random noise through the denoising module of the diffusion model, and the above denoising module is trained based on a plurality of first noise data and a plurality of second noise data. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition cost, therefore, in this solution, the acquisition cost of the training data of the data generation device is low.
[0038] Fifth aspect, the present application further provides an electronic device, including a processor and a memory. The processor and the memory are connected, and the memory is configured to store program code, and the processor is configured to call the program code to execute the method described in the first aspect or the second aspect.
[0039] Sixth aspect, the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method described in the first aspect or the second aspect.
[0040] Seventh aspect, the present application further provides a computer program product including instructions, when the computer program product runs on a computer, it causes the computer to execute the method described in the first aspect or the second aspect.
[0041] In an eighth aspect, the present application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface and executes the method described in the first aspect or the second aspect.
[0042] Optionally, as an implementation, the chip may further include a memory, in which instructions are stored, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the processor is configured to execute the method described in the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The following introduces the drawings used in the embodiments of the present application.
[0044] Figure 1 A schematic diagram of a diffusion model provided for an embodiment of the present application;
[0045] Figure 2A A schematic diagram of a model training method provided for an embodiment of the present application;
[0046] Figure 2B A schematic diagram of a system architecture provided for an embodiment of the present application;
[0047] Figure 3 A flowchart of a model training method provided for an embodiment of the present application;
[0048] Figure 4A A flowchart of another model training method provided for an embodiment of the present application;
[0049] Figure 4B A schematic diagram of another model training method provided for an embodiment of the present application;
[0050] Figure 5 A flowchart of a data generation method provided for an embodiment of the present application;
[0051] Figure 6 A flowchart of another data generation method provided for an embodiment of the present application;
[0052] Figure 7 A schematic diagram of the structure of a model training device provided for an embodiment of the present application;
[0053] Figure 8 A schematic diagram of the structure of a data generation device provided for an embodiment of the present application;
[0054] Figure 9 A schematic diagram of the structure of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0056] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0057] "At least one" mentioned in the embodiments of the present application means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural items (s). For example, at least one (item) of a, b, or c can represent: a, b, c, (a and b), (a and c), (b and c), or (a, b, and c), where a, b, c can be single or multiple. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. And the serial numbers of the steps in the embodiments of the present application (such as step S1, step S21, etc.) are only used to distinguish different steps and do not limit the order of execution between steps.
[0058] Moreover, unless otherwise stated, the ordinal numbers such as "first" and "second" used in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority, or importance of multiple objects. For example, the first device and the second device are only for ease of description and do not indicate differences in the structure, importance, etc. of the first device and the second device. In some embodiments, the first device and the second device can also be the same device.
[0059] As used in the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the concept and principles of the present application should be included in the protection scope of the present application.
[0060] Diffusion models are a class of generative models that learn the probability distribution of data by applying a series of random perturbations to the data. These perturbations are usually Gaussian noise, which gradually destroys the data until the data becomes pure noise. Refer to Figure 1 ,Figure 1 A schematic diagram of a diffusion model provided by an embodiment of the present application; the noise addition module of the diffusion model performs a series of noise addition processes (e.g., T times of noise addition processes) on the data X0 until the data X0 becomes pure noise Z. The above T is the maximum number of noise addition times, and the specific value can be set according to the actual situation. For example, T is 1000 or more than 1000. The generation process of the diffusion model is reversed by applying a series of denoising steps, and these steps gradually recover the original data X0 from the pure noise Z. Exemplarily, the denoising step is for a neural network to learn the conditional distribution of the next state, and this neural network is trained to minimize the reconstruction error.
[0061] The diffusion model can sample from the learned distribution to generate realistic and diverse data samples, such as images, texts, audios, and videos. The diffusion model can also generate data related to given conditions according to additional information, such as texts or audios. For example, generate an image that matches the description of a text prompt; generate a video that conforms to the description of a text prompt; generate a three-dimensional model that corresponds to the shape and appearance given in the text prompt description; generate an action sequence that matches the action or style given in the text prompt description; generate an image related to the given sound; generate a video synchronized with the given sound; generate dance movements harmonious with the given music.
[0062] Exemplarily, a task of generating human motion from a natural language description, such as "a person walks slowly and then runs fast" or "a person does a somersault". This task can be used in applications such as animation, simulation, and motion editing.
[0063] Exemplarily, a task of generating human motion from action labels, such as "walking", "jumping", or "dancing". This task can be used in applications such as motion synthesis, motion retrieval, and motion conversion.
[0064] Another example is a task of generating human motion from music, such as generating dance movements synchronized with the music rhythm and style. This task can be used in applications such as entertainment, education, and music analysis.
[0065] Training the neural network in the denoising step requires relying on high-quality labeled training data. However, the acquisition cost of training data with high quality and rich annotations is very high. For example, taking the training data as motion data, expensive motion capture systems and professional operators are needed to collect high-quality motion data, and annotating the motion data also requires a large amount of human resources and time.
[0066] Therefore, the embodiment of the present application provides a model training method, which can effectively reduce the acquisition cost of training data.
[0067] Refer to Figure 2A , Figure 2ASchematic diagram of a model training method provided by an embodiment of the present application; in the model training stage, the present application performs diffusion model training by combining unlabeled high-quality training data and labeled low-quality training data, so that the diffusion model can learn the features of high-quality training data and the label features of low-quality training data at the same time. In the model inference stage, only by inputting specified conditions and random noise, the trained diffusion model can output corresponding generated data. The present application can effectively utilize unlabeled high-quality training data and labeled low-quality training data to complete diffusion model training, which can not only reduce the acquisition cost of the training data of the diffusion model, but also ensure the quality of the generated data of the diffusion model.
[0068] The model training method of the embodiment of the present application is applicable to the generation tasks of various modal data, such as motion sequences, pictures, audio, video, 3D models, etc. The application scenarios of the generation task can include multimedia digital human dance generation, generating pictures, audio, video, etc. specified by users. The above-mentioned multimedia digital human dance generation refers to the process of using computer technology and artificial intelligence technology to create and generate digital human dances. Digital human dance generation can be widely applied in the fields of entertainment, game development, advertising, education, and art. For example, the video of digital human dance generation can be published on various online media and social media platforms. Digital human dance generation technology can be used to create digital human dance performances in virtual reality (VR) or augmented reality (AR) environments. Digital human dance generation technology can be used in the fields of advertising and marketing to create attractive visual effects and brand images. Brands can use digital human dance videos to showcase the features and advantages of products or services, attract potential customers, and improve brand awareness. Digital human dance generation technology can be used in the fields of education and training, such as dance education, sports training, etc. Digital human dances can be used as teaching demonstrations or training aids to help students or athletes better understand and master dance movements and skills. Digital human dance generation technology can not only create amazing visual effects, but also bring a brand-new viewing experience to the audience.
[0069] As long as the cost of high-quality training data with labels is high and it is difficult to obtain, while low-quality training data with labels and high-quality data without labels are relatively easy to obtain, the model training method of the embodiment of the present application can be used to reduce the acquisition cost of training data while ensuring the model effect.
[0070] The model training method of the embodiment of the present application can be executed by a model training device or by a chip in the model training device.
[0071] The following introduces a system architecture provided by an embodiment of the present application.
[0072] See the appendix Figure 2B As shown in the system architecture 200 provided by an embodiment of the present application, the data acquisition device 260 is used to acquire training data for the diffusion model. In this embodiment, the training data includes a plurality of first training data and a plurality of second training data. The data acquisition device 260 stores the training data in the database 230.
[0073] The training device 220 can perform model training based on the training data maintained in the database 230 to obtain the diffusion model 201. The training device 220 is the model training device in the embodiment of the present application. The specific model training process can refer to the following specific description of the model training method (such as Figure 3 ), and will not be elaborated here. Exemplarily, in the embodiment of the present application, after the model training is completed, the diffusion model 201 can obtain the ability to generate specified data based on specified conditions and random noise. The specific data generation can refer to the following specific description of the data generation method (such as Figure 5 ), and will not be elaborated here. The training device 220 can be a server, a cloud service device, etc., and can also be a mobile terminal, a tablet computer, a laptop computer, an AR / VR device, a vehicle-mounted terminal, a monitoring device, a vehicle-mounted autonomous driving system, a drone, etc.
[0074] It should be noted that in actual applications, the training data maintained in the database 230 may not all come from the acquisition of the data acquisition device 260, and may also be received from other devices. Additionally, it should be noted that the training device 220 may not necessarily perform model training completely based on the training data maintained in the database 230, and may also obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiment of the present application.
[0075] The diffusion model 201 obtained after being processed by the training device 220 can be applied and deployed in the system or device of the terminal, such as being applied to the Figure 2B shown terminal device 210. The terminal device 210 can process the input specified conditions and random noise based on the diffusion model 201 to obtain the corresponding model processing result (such as specified data). The terminal device 210 is the data generation device in the embodiment of the present application. The terminal device 210 can be a terminal, such as a mobile terminal, a tablet computer, a laptop computer, an AR / VR device, a vehicle-mounted terminal, a vehicle-mounted computing platform, a vehicle-mounted domain controller, a monitoring device, a vehicle-mounted autonomous driving system, a drone, etc., and can also be a server or a cloud device, etc. The training device 220 and the terminal device 210 can be the same device.
[0076] In the appendix Figure 2BAmong them, the terminal device 210 is configured with an I / O interface 212 for data interaction with external devices. The user can input data to the I / O interface 212 through the client device 240. In this embodiment, the input data is a specified condition or random noise, which can be input by the user or come from the database 230. Among them, the client device 240 can be an audio acquisition device, such as a mobile phone. For example, the mobile phone obtains a specified condition (in this case, the specified condition can be audio).
[0077] Optionally, the preprocessing module 213 is used to preprocess the input data received by the I / O interface 212. In the embodiment of the present application, the preprocessing module 213 is used to preprocess the input data received by the I / O interface 212, and the preprocessed data enters the calculation module 211. In the embodiment of the present application, the input data can be audio, and the preprocessing module 213 can be used to perform at least one preprocessing operation on the audio, such as feature extraction, sampling, framing, windowing, noise reduction, normalization, etc.
[0078] Among them, feature extraction is to extract features from each frame for subsequent classification or recognition. Sampling is to convert a continuous-time signal into a discrete-time signal for easy computer processing. Framing is to divide the speech signal into short segments, usually 20 - 40 milliseconds, to facilitate feature extraction. Windowing is to apply a window function, such as a Hamming window, to each frame to reduce spectral leakage. Noise reduction is to remove background noise in the audio signal to improve speech quality. And normalization is to normalize the amplitude of the audio signal to the same level, which helps with subsequent processing.
[0079] When the terminal device 210 preprocesses the input data, or when the calculation module 211 of the terminal device 210 performs calculations and other related processing processes, the terminal device 210 can call data, code, etc. in the data storage system 250 for corresponding processing, or can store the data, instructions, etc. obtained from the corresponding processing in the data storage system 250.
[0080] Finally, the I / O interface 212 returns the model processing result of the input data to the client device 240 and thus provides it to the user. At this time, the client device 240 can be a display.
[0081] Optionally, the model processing result can also then be used as the input of the calculation module, and the calculation module performs other processing operations according to the model processing result.
[0082] In the appendix Figure 2BIn the case shown, the user can manually input data, and this manual input can be operated through the interface provided by the I / O interface 212. In another case, the client device 240 can automatically send input data to the I / O interface 212. If the client device 240 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 240. The user can view the results output by the terminal device 210 on the client device 240, and the specific presentation forms can be display, sound, movement, etc. The client device 240 can also be used as a data acquisition end to collect, such as Figure 2B the input data input to the I / O interface 212 and the output results of the output I / O interface 212 shown as new sample data, and store them in the database 230. Of course, it can also be collected without going through the client device 240, but directly by the I / O interface 212, such as Figure 2B the input data input to the I / O interface 212 and the output results of the output I / O interface 212 shown as new sample data and store them in the database 230.
[0083] It should be noted that the appendix Figure 2B is only a schematic diagram of a system architecture provided by the embodiments of the present application. Figure 2B The positional relationships between the devices, components, modules, etc. shown in the appendix do not constitute any restrictions. For example, in the appendix Figure 2B the data storage system 250 is an external memory relative to the terminal device 210. In other cases, the data storage system 250 can also be placed in the terminal device 210.
[0084] The model training method of the embodiments of the present application will be specifically described below.
[0085] Referring to Figure 3 , Figure 3 which is a schematic flowchart of a model training method provided by the embodiments of the present application; in the embodiments of the present application, taking the model training device as an example, the model training method includes the following steps:
[0086] 301. The model training device obtains a plurality of first noise data corresponding to a plurality of first training data.
[0087] Specifically, the above-mentioned first training data has labels. The first noise data corresponding to the first training data is obtained by adding noise to the first training data for the first number of noise addition times. The first number of noise addition times is greater than the second number of noise addition times and less than or equal to the maximum number of noise addition times; the second number of noise addition times is related to the noise level of the plurality of first training data.
[0088] Among them, the specific form of the above-mentioned tag can be text and / or audio, etc. The tag of the above first training data can be understood as the specified condition or given condition when generating relevant data based on the first training data. The above text can be a sentence (such as "A person walks slowly and then runs fast") or a word (such as "jump", "run").
[0089] The above noise addition times can be understood as the noise addition step size, the noise addition time step size, or the degree of noise addition. For example, the larger the value of the first noise addition times, the longer the noise addition time for the first training data. On the contrary, the smaller the value of the first noise addition times, the less the noise addition time for the first training data. Also, for example, the larger the value of the first noise addition times, the stronger the noise perturbation applied to the first training data. On the contrary, the smaller the value of the first noise addition times, the weaker the noise perturbation applied to the first training data.
[0090] The above second noise addition times T* is the noise addition times determined based on the noise levels of multiple first training data.
[0091] The above maximum noise addition times T is the noise addition times required to turn the data into pure Gaussian noise, and the specific value can be set according to the actual situation. For example, the maximum noise addition times is 1000 or more than 1000, without limitation.
[0092] For the first noise addition times t1 of each first training data, it satisfies: t1 ∈ [T * +1, T]. Exemplarily, for each first training data, a value can be randomly selected from [T * +1, T] as the first noise addition times t1 of the first training data. For example, assume T* is 100 and T is 1000, and the randomly selected value for the first training data is 200. This indicates that the first training data needs to be subjected to 200 times of noise addition processing or the application of noise equivalent to 200 times of noise addition processing to obtain the first noise data corresponding to the first training data.
[0093] 302. The model training device obtains multiple second noise data corresponding to multiple second training data.
[0094] Specifically, the above second training data has no label, and the noise in the first training data is higher than the noise in the second training data. The second noise data corresponding to the second training data is obtained after adding noise to the second training data for the third noise addition times; the third noise addition times is less than or equal to the second noise addition times and greater than or equal to one.
[0095] The first training data and the second training data can be various modalities of data, such as pictures, audio, video, motion sequences, 3D models, etc.
[0096] Exemplarily, the first training data can be understood as labeled low-quality data, and the second training data can be understood as unlabeled high-quality data. The above quality refers to the quality of the data, that is, the noise contained in the data. The noise in the first training data is higher than that in the second training data. For example, the first training data is a motion sequence with text labels but inaccurate motion; such as a motion sequence extracted from labeled video data. The second training data is a motion sequence with accurate motion but without labels; such as the data of the Archive of MotionCapture As Surface Shapes (AMASS) human motion pose dataset. The data in this dataset is mainly collected by professional operators in a motion capture system but without additional annotation.
[0097] Further exemplarily, the noise of the first training data is greater than or equal to the noise threshold, while the noise of the second training data is less than the above noise threshold. The specific value of the noise threshold can be set according to the actual situation and is not limited. For example, when the training data is a picture, the noise can be understood as random changes in image information or pixel brightness. The first training data can be a labeled ordinary picture (obtained by an ordinary camera), and the second training data is an unlabeled high-definition picture (obtained by a high-definition camera). Another example is that when the training data is a motion sequence, the noise can be understood as interference such as jitter and ghosting. The first training data is a motion sequence with labels and jitter or ghosting, and the second training data is an unlabeled AMSS motion sequence.
[0098] In the embodiments of the present application, different types of training data adopt different sampling training ranges, that is, the number of noise addition times is different. For the third noise addition times t3 of each second training data, it satisfies: t3 ∈ [1, T * . Exemplarily, for each second training data, a value can be randomly selected from [1, T * as the third noise addition times t3 of this second training data. For example, assume T* is 100, and the randomly selected value for the second training data is 100, which means that the second training data needs to be subjected to 100 noise addition processes or noise equivalent to 100 noise addition processes to obtain the second noise data corresponding to the second training data.
[0099] The execution order of steps 301 and 302 is not limited. Steps 301 and 302 can be executed simultaneously, or the above two steps can be executed successively.
[0100] 303. The model training device trains the first denoising module of the diffusion model based on a plurality of first noise data and a plurality of second noise data to obtain a second denoising module.
[0101] In the embodiments of the present application, the first denoising module of the diffusion model is trained based on multiple first noise data and multiple second noise data to obtain a second denoising module. Since the first noise data and the second noise data are obtained based on the first training data and the second training data with low acquisition costs respectively, this solution can effectively reduce the acquisition costs of the training data of the diffusion model, including both the data acquisition cost and the data annotation cost.
[0102] In addition, the embodiments of the present application can effectively utilize the first training data and the second training data, enabling the second denoising module to learn both the label features of the first training data and the high-quality data features of the second training data simultaneously, and the second denoising module can achieve high-quality data generation. In the embodiments of the present application, even with only limited low-quality data, controllable and high-quality data can be generated. The data generation effect of the method in the embodiments of the present application reaches the same level as that of the denoising module trained with labeled high-quality training data.
[0103] In one possible implementation manner, refer to Figure 4A , Figure 4A which is a schematic flowchart of another model training method provided by the embodiments of the present application; the above model training method further includes the following steps:
[0104] 401. The model training device obtains multiple first training data and multiple second training data.
[0105] 402. The model training device performs noise addition processing on the multiple first training data through the noise addition module of the diffusion model to obtain multiple first noise data.
[0106] Specifically, after determining a first noise addition number for the first training data, the model training device performs noise addition for the first noise addition number on the first training data through the noise addition module of the diffusion model to obtain the corresponding first noise data. For example, when the first training data is a picture, after performing noise addition for the first noise addition number on the picture, a picture with noise, that is, the first noise data, can be obtained.
[0107] 403. The model training device performs noise addition processing on the multiple second training data through the noise addition module to obtain multiple second noise data.
[0108] Specifically, after determining a third noise addition number for the second training data, the model training device performs noise addition for the third noise addition number on the second training data through the noise addition module of the diffusion model to obtain the corresponding second noise data.
[0109] In the embodiments of the present application, the noise addition module of the diffusion model is used to obtain the first noise data and the second noise data.
[0110] Exemplarily, the above-mentioned second noise addition times T* is the number of noise addition times determined based on the noise levels of multiple first training data, and the second noise addition times T* can be determined through the following two examples.
[0111] In the first example, first determine the number of noise addition times corresponding to the noise values of multiple first training data, and then determine an average number of noise addition times based on the number of noise addition times corresponding to multiple first training data. This average number of noise addition times is the second noise addition times T*. The number of noise addition times corresponding to the noise value of the first training data can be obtained based on various methods. For example, assuming that the initial noise of the second training data is zero, the model training device performs multiple noise addition processes on the second training data to obtain the noise values of the second training data corresponding to different numbers of noise addition times; for example, the noise value corresponding to performing 100 noise additions on the second training data, the noise value corresponding to performing 500 noise additions on the second training data, or the noise value corresponding to performing 1000 noise additions on the second training data. Establish a first table based on the noise values of the second training data corresponding to different numbers of noise addition times. Determine the noise value of each first training data based on the noise evaluation equation, and then look up the first table according to the noise value of the first training data to determine the number of noise addition times corresponding to the first training data. The above noise evaluation equation is not particularly limited.
[0112] In the second example, select K first training data from multiple first training data, where K is greater than one, and the specific value of K can be determined according to the actual situation. First determine the average noise of the K first training data (which can be the average noise calculation based on the noise evaluation equation or the average noise evaluated by the human eye), and then perform noise addition on K second training data among multiple second training data until the average noise of the K second training data is close to the above average noise. Then the number of noise addition times of the second training data is the second noise addition times T*. The average noise of the above K second training data can be the average noise calculation based on the noise evaluation equation or the average noise evaluated by the human eye.
[0113] In a possible implementation manner, the above step 303 specifically includes the following steps:
[0114] The model training device trains the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data, to obtain a second denoising module.
[0115] Specifically, the model training device simultaneously conducts training and learning on the first training data and the second training data. For the first training data, the model training device conducts model training based on the label of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data. At this time, the model focuses on learning the label features of the first training data. For the second training data, the model training device conducts model training based on the second noise data and the third noise addition times corresponding to the second noise data. At this time, the model focuses on learning the high-quality data features of the second training data.
[0116] In the embodiments of the present application, the second denoising module can simultaneously learn the label features of the first training data and the high-quality data features of the second training data, enabling the second denoising module to achieve high-quality data generation and with low training costs.
[0117] Further exemplarily, the above model training device trains the first denoising module based on the labels of multiple first training data, multiple first noise data, and multiple first noise addition times corresponding to the multiple first noise data, and based on multiple second noise data and multiple third noise addition times corresponding to the multiple second noise data to obtain the second denoising module, which specifically includes the following steps:
[0118] S1. The model training device predicts the first noise value of the first noise data through the first denoising module based on the label of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data.
[0119] The above first noise value is a predicted value of the noise addition amount of the Nth noise addition to the first training data, where N is the first noise addition times.
[0120] Specifically, refer to Figure 4B , Figure 4B which is a schematic diagram of another model training method provided in the embodiments of the present application; among them, by inputting the first training data X' and the second training data X into the diffusion model, the first noise data corresponding to the first training data X' and the second noise data corresponding to the second training data X can be obtained first. c represents the label of the first training data X', is an empty set, used to improve the robustness of the diffusion model; a represents a word, w represents a sentence, and m represents an audio. That is, the label of the first training data can be a single word, or a sentence or an audio. t1 is the first noise addition times corresponding to the first training data X', and t3 is the third noise addition times corresponding to the second training data X.
[0121] The diffusion model predicts the first noise value of the first noise data through the first denoising module based on the label of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data
[0122] S2. The model training device predicts the second noise value of the second noise data through the first denoising module based on the second noise data and the third noise addition times corresponding to the second noise data.
[0123] The above second noise value is a predicted value of the noise addition amount for the M-th noise addition to the second training data, where M is the third noise addition times.
[0124] Specifically, referring to Figure 4B , the diffusion model predicts the second noise value of the second noise data through the first denoising module based on the second noise data and the third noise addition times corresponding to the second noise data
[0125] S3. The model training device adjusts the first denoising module based on the first loss value and the second loss value.
[0126] The above first loss value is obtained based on the first noise value and the third noise value, and the third noise value is the true value of the noise addition amount for the N-th noise addition. The second loss value is obtained based on the second noise value and the fourth noise value, and the fourth noise value is the true value of the noise addition amount for the M-th noise addition.
[0127] Specifically, referring to Figure 4B , the first noise value is supervised by the third noise value Sx', and based on the first noise value and the third noise value Sx', the first loss value can be obtained. The second noise value is supervised by the fourth noise value Sx, and based on the second noise value and the fourth noise value Sx, the second loss value can be obtained. The model training device adjusts the first denoising module based on the first loss value and the second loss value.
[0128] In the embodiment of the present application, when training the first denoising module, by predicting the first noise value and the second noise value, the first loss value and the second loss value can be obtained, and then the first denoising module can be adjusted based on the first loss value and the second loss value. Repeat the execution until the training stop condition is satisfied, and then the above second denoising module can be obtained. The above training stop condition can be set according to the actual situation. For example, the above training stop condition is that the third loss value satisfies a certain condition (such as the third loss value is lower than a certain threshold), or the change of the third loss value in multiple consecutive times is very small (such as the change of the third loss value in several consecutive batches is very small, and each batch can be 5 or 10 third loss values).
[0129] In a possible implementation manner, the above step S3 specifically includes the following steps:
[0130] The model training device determines the third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value.
[0131] Specifically, the specific values of the first weight value and the second weight value can be set according to the actual situation without limitation. For example, the first weight value and the second weight value can be the same, for example, both are 0.5, which means that the model learning pays equal attention to the label features and the high-quality training data features. For another example, the first weight value is 0.6 and the second weight value is 0.4, indicating that the model learning pays more attention to learning the label features. For another example, the first weight value is 0.4 and the second weight value is 0.6, indicating that the model learning pays more attention to learning the high-quality training data features.
[0132] In the embodiments of the present application, the model training device can set the first weight value and the second weight value to adjust the focus of model learning to meet different model training requirements.
[0133] The present application also provides a data generation method, which can be executed by a data generation device or by a chip in the data generation device.
[0134] Refer to Figure 5 , Figure 5 which is a schematic flow chart of a data generation method provided by the embodiments of the present application; taking the data generation device as an example of the execution subject, the above data generation method includes the following steps:
[0135] 501. The data generation device obtains indication information and random noise.
[0136] The above indication information is related to the first data, that is, the indication information is used to indicate the generation object or generation target of the denoising module, that is, the given condition.
[0137] Exemplarily, sampling a noise from a standard Gaussian distribution can obtain the above random noise.
[0138] 502. The data generation device generates the first data through the denoising module of the diffusion model based on the indication information and the random noise.
[0139] The above denoising module is trained by using the above model training method.
[0140] The generated first data can be various modalities of data, such as pictures, audio, video, motion sequences, 3D models, etc. The specific modality of the first data is the same as that of the first training data and the second training data. For example, when the first training data and the second training data are pictures, the first data is a picture. For another example, when the first training data and the second training data are motion sequences, the first data is a motion sequence.
[0141] In the embodiments of the present application, the denoising module can generate first data based on the indication information and random noise. The above denoising module is trained based on a plurality of first noise data and a plurality of second noise data. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, in this solution, the acquisition cost of the training data of the denoising module is low.
[0142] Reference Figure 6 , Figure 6 is a schematic flowchart of another data generation method provided by the embodiments of the present application; assume that the indication information is c and the random noise is X T ∈N(0, I). Among them, N(0, I) is a normal distribution with a mean of 0 and a variance of 1, and I represents the variance. That is, the random noise satisfies the normal distribution (Gaussian distribution). After completing the training of the denoising module of the diffusion model using the above model training method, the denoising module knows T * . During the data generation process, before T * , the denoising module mainly generates approximate data based on the guidance of the indication information, that is, sub-high-quality data; after T * , the denoising module supplements high-quality details based on the sub-high-quality data to obtain the final first data X0. The result of this method is basically the same as the result of the labeled high-quality data set.
[0143] The above elaborates in detail the method of the embodiments of the present application. Next, the device provided by the embodiments of the present application will be introduced.
[0144] Figure 7 , Figure 8 and Figure 9 are schematic structural diagrams of possible devices provided by the embodiments of the present application. Among them, Figure 7 The model training device shown can be used to implement the functions of the above Figure 3 shown model training method embodiment, and thus can also achieve the beneficial effects possessed by the above model training method embodiment. In the embodiments of the present application, the model training device can be an electronic device or a module (such as a chip) applied to an electronic device. Figure 8 The data generation device shown can be used to implement the functions of the above Figure 5 shown data generation method embodiment, and thus can also achieve the beneficial effects possessed by the above data generation method embodiment. In the embodiments of the present application, the data generation device can be an electronic device or a module (such as a chip) applied to an electronic device.
[0145] As Figure 7 shown, Figure 7Schematic structural diagram of a model training device provided by an embodiment of the present application; the model training device 700 includes an acquisition module 710 and a training module 720. The model training device 700 is used to implement the above Figure 3 functions of the model training method embodiment shown in. Alternatively, the model training device 700 may include a module for implementing any function or operation of the model training method embodiment shown in the above Figure 3 , and this module can be implemented in whole or in part by software, hardware, firmware, or any combination thereof.
[0146] When the model training device 700 is used to implement the Figure 3 functions of the method embodiment shown, the acquisition module 710 is used to acquire multiple first noise data corresponding to multiple first training data. The above first training data has labels, and the first noise data corresponding to the first training data is obtained by adding noise to the first training data for the first noise addition times. The first noise addition times is greater than the second noise addition times, and the first noise addition times is less than or equal to the maximum noise addition times; the second noise addition times is related to the noise levels of the multiple first training data. The acquisition module 710 is further used to acquire multiple second noise data corresponding to multiple second training data. The above second training data has no labels, and the noise in the first training data is higher than the noise in the second training data. The second noise data corresponding to the second training data is obtained by adding noise to the second training data for the third noise addition times. The third noise addition times is less than or equal to the second noise addition times, and the third noise addition times is greater than or equal to one. The training module 720 is used to train the first denoising module of the diffusion model based on the multiple first noise data and the multiple second noise data to obtain a second denoising module.
[0147] In the embodiment of the present application, the model training device 700 trains the first denoising module of the diffusion model based on the multiple first noise data and the multiple second noise data to obtain a second denoising module. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, the acquisition cost of the training data of the model training device in this solution is low.
[0148] In a possible implementation manner, the above training module 720 is specifically used for:
[0149] Based on the labels of the multiple first training data, the multiple first noise data, and the multiple first noise addition times corresponding to the multiple first noise data, and based on the multiple second noise data and the multiple third noise addition times corresponding to the multiple second noise data, train the first denoising module to obtain a second denoising module.
[0150] In a possible implementation, the above-mentioned training module 720 is specifically used for training the first denoising module based on the labels of a plurality of first training data, a plurality of first noise data, and a plurality of first noise addition times corresponding to the plurality of first noise data, and based on the plurality of second noise data and a plurality of third noise addition times corresponding to the plurality of second noise data to obtain a second denoising module, specifically as follows:
[0151] Based on the labels of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data, predict the first noise value of the first noise data through the first denoising module. The above first noise value is the predicted value of the noise addition amount for the Nth noise addition, where N is the first noise addition times.
[0152] Based on the second noise data and the third noise addition times corresponding to the second noise data, predict the second noise value of the second noise data through the first denoising module. The above second noise value is the predicted value of the noise addition amount for the Mth noise addition, where M is the third noise addition times.
[0153] Adjust the first denoising module based on the first loss value and the second loss value. The first loss value is obtained based on the first noise value and the third noise value, and the third noise value is the true value of the noise addition amount for the Nth noise addition. The second loss value is obtained based on the second noise value and the fourth noise value, and the fourth noise value is the true value of the noise addition amount for the Mth noise addition.
[0154] In a possible implementation, the above-mentioned training module 720 is specifically used for adjusting the first denoising module based on the first loss value and the second loss value, specifically as follows:
[0155] Determine the third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value. Adjust the first denoising module based on the third loss value.
[0156] In a possible implementation, the above-mentioned acquisition module 710 is further used to acquire a plurality of first training data and a plurality of second training data.
[0157] The above-mentioned model training device 700 further includes a processing module 730.
[0158] The processing module 730 is used to perform noise addition processing on a plurality of first training data through the noise addition module of the diffusion model to obtain a plurality of first noise data.
[0159] The above-mentioned processing module 730 is further used to perform noise addition processing on a plurality of second training data through the noise addition module to obtain a plurality of second noise data.
[0160] For the introduction of the above-mentioned modules, reference can be made to the descriptions in the foregoing embodiments, which will not be elaborated herein.
[0161] As Figure 8 shown,Figure 8 Schematic structural diagram of a data generation device provided in an embodiment of the present application; the data generation device 800 includes an acquisition module 810 and a generation module 820. The data generation device 800 is used to implement the above Figure 5 functions of the data generation method embodiments shown. Alternatively, the data generation device 800 may include modules for implementing any function or operation of the data generation method embodiments shown above Figure 5 , and this module can be implemented in whole or in part by software, hardware, firmware, or any combination thereof.
[0162] When the data generation device 800 is used to implement the Figure 5 functions of the method embodiments shown, the acquisition module 810 is used to acquire indication information and random noise. The indication information is related to the first data. The generation module 820 is used to generate the first data through the denoising module of the diffusion model based on the indication information and random noise. The above denoising module is trained by using the model training method described in the first aspect.
[0163] In the embodiments of the present application, the data generation device can generate the first data through the denoising module of the diffusion model based on the indication information and random noise, and the above denoising module is trained based on a plurality of first noise data and a plurality of second noise data. Since the first noise data and the second noise data are respectively obtained based on the first training data and the second training data with low acquisition costs, in this solution, the acquisition cost of the training data of the data generation device is low.
[0164] For the introduction of the above modules, reference can be made to the records of the foregoing embodiments, which will not be elaborated here.
[0165] Refer to Figure 9 , Figure 9 Schematic structural diagram of an electronic device provided in an embodiment of the present application; the electronic device 900 includes a memory 901, a processor 902, a communication interface 904, and a bus 903. Among them, the memory 901, the processor 902, and the communication interface 904 are communicatively connected to each other through the bus 903.
[0166] Optionally, the above electronic device 900 further includes a display screen (not shown), and the display screen is communicatively connected to the memory 901, the processor 902, and the communication interface 904 through the bus 903. The display screen is used to output information for interaction with the user, such as voice output or display output.
[0167] The memory 901 can be a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM). The memory 901 can store a program. When the program stored in the memory 901 is executed by the processor 902, the processor 902 and the communication interface 904 are used to execute each step of the model training method or the data generation method of any embodiment of this application.
[0168] The processor 902 can be a general-purpose Central Processing Unit (CPU), a microprocessor, an Application Specific Integrated Circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the functions required to be executed by the units in the model training device or the data generation device of any embodiment of this application, or to execute the model training method or the data generation method of any embodiment of this application.
[0169] The processor 902 can also be an integrated circuit chip with signal processing capabilities. During implementation, each step of the model training method or the data generation method of any embodiment of this application can be completed by the integrated logic circuit in the hardware of the processor 902 or by instructions in software form. The above-mentioned processor 902 can also be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the model training method, the data generation method, the steps, and the logic block diagram disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The steps of the model training method or the data generation method of any embodiment of this application in combination can be directly embodied as being completed by the execution of the hardware processor, or by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 901, and the processor 902 reads the information in the memory 901 and combines its hardware to complete the functions required to be executed by the units included in the model training device or the data generation device of any embodiment of this application, or to execute the model training method or the data generation method of any embodiment of this application.
[0170] The communication interface 904 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the electronic device 900 and other devices or communication networks. For example, indication information and the like can be obtained through the communication interface 904.
[0171] The bus 903 may include a path for transmitting information between various components of the electronic device 900 (for example, the memory 901, the processor 902, the communication interface 904). In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0172] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0173] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0174] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape, a magnetic disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state disk (SSD), etc.
[0175] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A model training method, characterized in that, The method includes: Obtaining a plurality of first noise data corresponding to a plurality of first training data; the first training data has labels, and the first noise data corresponding to the first training data is obtained by adding noise to the first training data for a first number of noise addition times; the first number of noise addition times is greater than a second number of noise addition times, and the first number of noise addition times is less than or equal to a maximum number of noise addition times; the second number of noise addition times is related to the noise level of the plurality of first training data; Obtaining a plurality of second noise data corresponding to a plurality of second training data; the second training data has no labels, and the noise in the first training data is higher than the noise in the second training data; the second noise data corresponding to the second training data is obtained by adding noise to the second training data for a third number of noise addition times; the third number of noise addition times is less than or equal to the second number of noise addition times, and the third number of noise addition times is greater than or equal to one; Training a first denoising module of a diffusion model based on the plurality of first noise data and the plurality of second noise data to obtain a second denoising module.
2. The method according to claim 1, wherein The training of the first denoising module of the diffusion model based on the plurality of first noise data and the plurality of second noise data to obtain a second denoising module includes: Training the first denoising module based on the labels of the plurality of first training data, the plurality of first noise data, and the plurality of first noise addition times corresponding to the plurality of first noise data, and based on the plurality of second noise data and the plurality of third noise addition times corresponding to the plurality of second noise data, to obtain the second denoising module.
3. The method according to claim 2, wherein The training of the first denoising module based on the labels of the plurality of first training data, the plurality of first noise data, and the plurality of first noise addition times corresponding to the plurality of first noise data, and based on the plurality of second noise data and the plurality of third noise addition times corresponding to the plurality of second noise data, to obtain the second denoising module includes: Based on the label of the first training data, the first noise data, and the first noise addition time corresponding to the first noise data, predicting a first noise value of the first noise data through the first denoising module; the first noise value is a predicted value of the noise addition amount of the Nth noise addition, where N is the first number of noise addition times; Based on the second noise data and the third noise addition time corresponding to the second noise data, predicting a second noise value of the second noise data through the first denoising module; the second noise value is a predicted value of the noise addition amount of the Mth noise addition, where M is the third number of noise addition times; Adjusting the first denoising module based on a first loss value and a second loss value; the first loss value is obtained based on the first noise value and a third noise value, the third noise value being the true value of the noise addition amount of the Nth noise addition, and the second loss value is obtained based on the second noise value and a fourth noise value, the fourth noise value being the true value of the noise addition amount of the Mth noise addition.
4. The method according to claim 3, wherein The adjusting the first denoising module based on the first loss value and the second loss value includes: Determine a third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value; Adjust the first denoising module based on the third loss value.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain the plurality of first training data and the plurality of second training data; Perform noise addition processing on the plurality of first training data through the noise addition module of the diffusion model to obtain the plurality of first noise data; Perform noise addition processing on the plurality of second training data through the noise addition module to obtain the plurality of second noise data.
6. A data generation method, characterized in that, The method includes: Obtain indication information and random noise; the indication information is related to first data; Generate the first data through the denoising module of the diffusion model based on the indication information and the random noise; the denoising module is trained by using the model training method according to any one of claims 1 to 5.
7. A model training device, characterized in that, Includes: An acquisition module for acquiring a plurality of first noise data corresponding to a plurality of first training data; The first training data is labeled, and the first noise data corresponding to the first training data is obtained after noise addition for a first number of noise addition times to the first training data; the first number of noise addition times is greater than a second number of noise addition times, and the first number of noise addition times is less than or equal to a maximum number of noise addition times; the second number of noise addition times is related to the noise level of the plurality of first training data; The acquisition module is further configured to acquire a plurality of second noise data corresponding to a plurality of second training data; the second training data is unlabeled, and the noise in the first training data is higher than the noise in the second training data; the second noise data corresponding to the second training data is obtained after noise addition for a third number of noise addition times to the second training data; the third number of noise addition times is less than or equal to the second number of noise addition times, and the third number of noise addition times is greater than or equal to one; A training module for training the first denoising module of the diffusion model based on the plurality of first noise data and the plurality of second noise data to obtain a second denoising module.
8. The device according to claim 7, characterized in that, Specifically, the training module is configured to: Train the first denoising module based on the labels of the plurality of first training data, the plurality of first noise data, and the plurality of first numbers of noise addition times corresponding to the plurality of first noise data, and based on the plurality of second noise data and the plurality of third numbers of noise addition times corresponding to the plurality of second noise data, to obtain the second denoising module.
9. The device according to claim 8, wherein Specifically, when the training module trains the first denoising module based on the labels of the plurality of first training data, the plurality of first noise data, and the plurality of first numbers of noise addition times corresponding to the plurality of first noise data, and based on the plurality of second noise data and the plurality of third numbers of noise addition times corresponding to the plurality of second noise data, to obtain the second denoising module, it is configured to: Predict a first noise value of the first noise data through the first denoising module based on the label of the first training data, the first noise data, and the first noise addition times corresponding to the first noise data; the first noise value is a predicted value of the noise addition amount for the Nth noise addition, and N is the first noise addition times. Predict a second noise value of the second noise data through the first denoising module based on the second noise data and the third noise addition times corresponding to the second noise data; the second noise value is a predicted value of the noise addition amount for the Mth noise addition, and M is the third noise addition times. Adjust the first denoising module based on a first loss value and a second loss value; the first loss value is obtained based on the first noise value and a third noise value, and the third noise value is the true value of the noise addition amount for the Nth noise addition, and the second loss value is obtained based on the second noise value and a fourth noise value, and the fourth noise value is the true value of the noise addition amount for the Mth noise addition.
10. The device according to claim 9, characterized in that, In terms of adjusting the first denoising module based on the first loss value and the second loss value, the training module is specifically configured to: Determine a third loss value based on the first loss value and its first weight value, and the second loss value and its second weight value; Adjust the first denoising module based on the third loss value.
11. The device according to any one of claims 7 to 10, wherein The obtaining module is further configured to obtain the plurality of first training data and the plurality of second training data; The device further comprises: A processing module, configured to perform noise addition processing on the plurality of first training data through the noise addition module of the diffusion model to obtain the plurality of first noise data; The processing module is further configured to perform noise addition processing on the plurality of second training data through the noise addition module to obtain the plurality of second noise data.
12. A data generation device, characterized in that, Comprises: An obtaining module, configured to obtain indication information and random noise; the indication information is related to first data; A generating module, configured to generate the first data through the denoising module of the diffusion model based on the indication information and the random noise; the denoising module is trained by using the model training method according to any one of claims 1 to 5.
13. An electronic device, characterized in that, Comprises a processor and a memory, wherein the processor is connected to the memory, and the memory is configured to store program code, and the processor is configured to call the program code to execute the method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 6.