Construction method and device of generative model, equipment and medium

By introducing a customized conditional injection network into the generative model and adopting a two-stage training method, the shortcomings of the existing generative model in terms of feature presentation effects are solved, and the feature extraction ability and presentation effect of the model are significantly improved.

CN120197652APending Publication Date: 2025-06-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311786637.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing generative models are not good in terms of feature presentation, especially in portrait generation tasks, which cannot accurately restore features such as facial facial features and skin types.

Method used

Based on the preset general generation network, the initial custom generation model is constructed by injecting the network into the network with customized conditions, and the model parameters are adjusted through the second-stage training method to ensure that parameter changes at different stages help to improve feature extraction capabilities.

Benefits of technology

By introducing custom conditional injection network and two-stage training methods, the feature extraction ability and feature presentation effect of the generative model are significantly improved, effectively solving the shortcomings of existing models in feature presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197652A_ABST
    Figure CN120197652A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a construction method and device of a generative model, equipment and a medium. The method comprises the steps that a preset universal generative network is acquired; constructing an initial custom generation model based on the general generation network and the custom condition injection network; performing first-stage training on the initial custom generation model to obtain a custom generation model after the first-stage training; parameters of a first feature extraction unit in the customized conditional injection network are changed during a first-stage training period; performing second-stage training on the custom generation model after the first-stage training to obtain a custom generation model after the second-stage training; partial network parameters of the general generation network and parameters of a second feature extraction unit in the customized condition injection network are changed during the two-stage training period, and parameters of the first feature extraction unit are not changed during the two-stage training period. According to the embodiment of the invention, the training difficulty can be effectively reduced, and the feature presentation effect of the generated model obtained by training is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, apparatus, device, and medium for constructing a generation model. Background Art

[0002] With the popularization of generation models, generation models are adopted in various fields such as the game field and the academic field to conveniently and quickly generate the required multimedia information. For example, an image generation model can be used to generate realistic or diverse images. However, the feature presentation effect of existing generation models is not good. Summary of the Invention

[0003] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for constructing a generation model.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for constructing a generation model. The method includes: obtaining a preset general generation network; constructing an initial custom generation model based on the general generation network and a custom conditional injection network; performing a first-stage training on the initial custom generation model to obtain a custom generation model after the first-stage training; wherein, parameters of a first feature extraction unit in the custom conditional injection network change during the first-stage training; performing a second-stage training on the custom generation model after the first-stage training to obtain a custom generation model after the second-stage training; wherein, some network parameters of the general generation network and parameters of a second feature extraction unit in the custom conditional injection network change during the second-stage training, parameters of the first feature extraction unit remain unchanged during the second-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0005] An embodiment of the present disclosure also provides an apparatus for constructing a generation model, including: a network acquisition module configured to acquire a preset general generation network; a model construction module configured to construct an initial custom generation model based on the general generation network and a custom conditional injection network; a first-stage training module configured to perform first-stage training on the initial custom generation model to obtain a custom generation model after the first-stage training; wherein, parameters of a first feature extraction unit in the custom conditional injection network are changed during the first-stage training; a second-stage training module configured to perform second-stage training on the custom generation model after the first-stage training to obtain a custom generation model after the second-stage training; wherein, some network parameters of the general generation network and parameters of a second feature extraction unit in the custom conditional injection network are changed during the second-stage training, parameters of the first feature extraction unit remain unchanged during the second-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0006] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for constructing a generation model provided by the embodiment of the present disclosure.

[0007] An embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program for executing the method for constructing a generation model provided by the embodiment of the present disclosure.

[0008] The above technical solution provided by the embodiment of the present disclosure can efficiently construct an initial custom generation model by combining a custom conditional injection network on the basis of a preset general generation network. By adopting a two-stage training method for the constructed generation model, the model parameters in different stages are different, which can effectively reduce the training difficulty. Moreover, since the constructed custom generation model introduces a custom conditional injection network on the basis of the general generation network, the feature extraction ability of the generation model can be better improved according to requirements, effectively ensuring the feature presentation effect of the trained generation model.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0010] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0011] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 It is a schematic flowchart of a method for constructing a generation model provided by an embodiment of the present disclosure;

[0013] Figure 2 It is a schematic structural diagram of a custom generation model provided by an embodiment of the present disclosure;

[0014] Figure 3 It is a schematic structural diagram of a custom generation model provided by an embodiment of the present disclosure;

[0015] Figure 4 It is a schematic structural diagram of a device for constructing a generation model provided by an embodiment of the present disclosure;

[0016] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specific embodiments

[0017] In order to better understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0018] Many specific details are set forth in the following description in order to provide a thorough understanding of the present disclosure, but the present disclosure may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0019] Figure 1 It is a schematic flowchart of a method for constructing a generation model provided by an embodiment of the present disclosure. This method can be executed by a device for constructing a generation model, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method mainly includes the following steps S102 to S108:

[0020] Step S102, obtain a preset general generation network.

[0021] The general generation network can be a pre-trained generation network, such as an open-source generation network, which includes the network structure required for the generation network. Exemplarily, the general generation network can adopt an existing general diffusion model. In some implementation examples, the general generation network includes an encoding network, a denoising network, and a decoding network connected in sequence. The encoding network and the decoding network are not limited in the embodiments of the present disclosure. Exemplarily, the encoding network and the decoding network can be implemented by the encoder and decoder in the VAE (Variational Auto-Encoder) model. The denoising network includes but is not limited to the Unet network. Exemplarily, in the image-to-image scenario, the general generation network can process the input image to generate a stylized image. However, the existing general generation network usually cannot well restore the detailed features of the input image. For example, the existing general generation network cannot accurately restore features such as the facial feature proportions and skin texture in the portrait generation task.

[0022] Step S104, based on the general generation network and the custom conditional injection network, construct an initial custom generation model. The structure of the conditional injection network is not limited in the embodiments of the present disclosure. For example, it can include functional units such as a feature extraction unit. The conditional injection network can not only extract features from the received data, but also perform processing such as feature mapping (this is only an example here and should not be regarded as a limitation. The specific processing operations to be performed can be flexibly set according to requirements), and inject the obtained conditional information such as feature-related information into the generation process. Taking the custom generation model to be finally constructed as an image generation model as an example, the conditional injection network in the custom generation model can, to a certain extent, guide image generation through processing operations such as feature extraction and information injection. The conditional injection network is connected to the general generation network. Different inputs of the model can be correspondingly input into the conditional injection network and the general generation network. The output result of the conditional injection network can be fused with the intermediate information generated by the general generation network, and the final generation result of the general generation network is used as the output result of the custom generation model. For ease of understanding, reference can be made to Figure 2 the structural schematic diagram of a custom generation model shown in Figure 2 In it, it is assumed that the input of the general generation network is the first input data, and the input of the custom conditional injection network is the second input data. The conditional injection network can perform feature extraction processing on the second input data, and perform fusion processing based on the extracted features and the intermediate data generated by the general generation network for the first input data, etc. Finally, the general generation network outputs the generated data. The types of the input data and the generated data are not limited in the embodiments of the present disclosure. For example, the first input data can be an image, the second input data can include text and images, and the generated data can be an image.

[0023] Step S106, perform a first-stage training on the initial custom generation model to obtain the custom generation model after the first-stage training; the parameters of the first feature extraction unit in the custom conditional injection network change during the first-stage training. Exemplarily, the first feature extraction unit may include an image encoder (Img-encoder). In some implementation examples, the parameters of the general generation network remain unchanged during the first-stage training. In addition, some parameters of the general generation network can also be changed during the first-stage training according to requirements, which is not limited herein.

[0024] In addition, it should be noted that in practical applications, the custom generation model may also include an Adapter unit. The Adapter unit can be implemented by networks such as MLP (Multilayer Perceptron), and specific references can be made to related technologies, which are not limited herein. The Adapter unit can map the features of the input data (such as images and / or texts) of the conditional injection network to the corresponding feature space, and the mapped features can be injected into the corresponding generation process (such as the image generation process) of the custom generation model through some networks in the custom generation model. The above-mentioned some networks can be, for example, the cross-attention unit of the denoising network (such as Unet) in the general generation network and the Lora layer added in the custom generation model. In this way, the conditional injection network can inject conditional information such as feature-related information into the generation process through the Adapter unit, so that the finally generated data (such as generated images) can reflect at least part of the features of the input data to a certain extent. During the training stage, the relevant network parameters of the Adapter unit will also be adjusted to have the expected information processing ability.

[0025] Step S108, perform a second-stage training on the custom generation model after the first-stage training to obtain the custom generation model after the second-stage training; wherein, some network parameters of the general generation network and the parameters of the second feature extraction unit in the custom conditional injection network change during the second-stage training, and the parameters of the first feature extraction unit remain unchanged during the second-stage training. The first feature extraction unit and the second feature extraction unit are used to extract different types of features, and the feature types include but are not limited to text features and image features. In some specific implementation examples, the first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder (Text-encoder).

[0026] In some specific implementation examples, some network parameters of the general generation network include the parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during both the first-stage training and the second-stage training, and the parameters of the first feature extraction unit remain unchanged during the second-stage training. Exemplarily, when the denoising network is a Unet network, the parameters of the Unet network only change during the second-stage training. Additionally, the second-stage training is also used to adjust the parameters of the second feature extraction unit. For example, the parameters of the text encoder are changed during the second-stage training, but the parameters of the image encoder have been adjusted in the first stage and remain unchanged during the second-stage training.

[0027] To ensure the training effect of the second stage, in the embodiments of the present disclosure, the LoRA technology (also known as LoRA fine-tuning technology or LoRA training technology) can be used to perform second-stage training on the custom generation model after the first-stage training. For example, the parameters of the denoising network (such as the Unet network) and the text encoder are adjusted based on the LoRA technology. The LoRA technology can assist model training by inserting some network layers (which can be called LoRA layers) into the model to be trained. During the second-stage training, the network parameters involved in the LoRA technology will also be adjusted accordingly. For specific details, reference can be made to the related technology and will not be elaborated here.

[0028] The above technical solutions provided by the embodiments of the present disclosure can efficiently construct an initial custom generation model by combining a custom conditional injection network on the basis of a preset general generation network. By adopting a two-stage training method for the constructed generation model, with different model parameters in different training stages, the training difficulty can be effectively reduced. Moreover, since the constructed custom generation model introduces a custom conditional injection network on the basis of the general generation network, the feature extraction ability of the generation model can be improved well according to requirements, effectively ensuring the feature presentation effect of the trained generation model.

[0029] In some implementation examples, the first-stage training of the initial custom generation model includes: using the preset initial training samples to perform first-stage training on the initial custom generation model; on this basis, the second-stage training of the custom generation model after the first-stage training includes: cleaning the initial training samples to obtain target training samples; using the target training samples to perform second-stage training on the custom generation model after the first-stage training. Compared with the related technology that directly continues to use the training samples adopted in the first-stage training for the second-stage training, the embodiments of the present disclosure can clean the initial training samples to obtain high-quality training samples more suitable for the second-stage training, thereby improving the training effect of the second stage.

[0030] Embodiments of the present disclosure provide an implementation example of cleaning an initial training sample to obtain a target training sample, which can be executed according to the following steps A to C:

[0031] Step A: Obtain the target information corresponding to the initial training sample; the target information includes the resolution of the significant object in the initial training sample and / or the clarity of the initial training sample. In practical applications, if the target information includes the resolution of the significant object, the significant object detection algorithm (or significant detection model) can be first used to detect the significant object in the initial training sample. The significant object can be, for example, a significant object such as a person in an image. Embodiments of the present disclosure do not limit the type of the significant object.

[0032] Step B: Based on the target information, perform a filtering operation on the initial training sample. Specifically, the initial training samples with poor target information can be filtered out, so as to retain high-quality samples.

[0033] Exemplarily, when the target information includes the resolution of the significant object in the initial training sample, the initial training samples with the resolution of the significant object less than the preset resolution threshold can be filtered out. For example, the preset resolution threshold can be 100*100 pixels. In this way, it can be effectively ensured that the significant objects in the remaining sample images have a proportion that meets the requirements.

[0034] When the target information includes the clarity of the initial training sample, the clarity threshold can be determined based on the clarity corresponding to each initial training sample; then the initial training samples with the clarity less than the clarity threshold are filtered out.

[0035] In some specific implementation examples, the clarity corresponding to the initial training sample is determined based on the Gaussian covariance value of the initial training sample; determining the clarity threshold based on the clarity corresponding to each initial training sample includes: performing an averaging process on the Gaussian covariance values corresponding to each initial training sample to obtain an average value; determining the clarity threshold based on the product of the average value and the preset ratio. The preset ratio can be flexibly set according to requirements, such as being set to 10%. In the above manner, it can be effectively ensured that the remaining sample images have high clarity.

[0036] Step C: Based on the initial training samples remaining after the filtering operation, obtain the target training sample. In some implementation examples, the initial training samples remaining after the filtering operation can be directly used as the target training sample. In other implementation examples, further operations can be performed on the initial training samples remaining after the filtering operation, such as cropping the main area therefrom to obtain the target training sample that meets the requirements. Exemplarily, it can be executed according to the following steps C1 to C3:

[0037] Step C1: Obtain the bounding rectangle of the salient object in the remaining initial training samples after the filtering operation. The bounding rectangle of the salient object can be obtained based on a salient object detection algorithm.

[0038] Step C2: Perform an expansion process on the bounding rectangle to obtain an expanded rectangle. In practical applications, the bounding rectangle can be expanded until its area or size is expanded by N times (N is a positive integer, such as 3 times), or until it is expanded to the boundary of the initial training samples.

[0039] Step C3: Crop the target region from the remaining initial training samples after the filtering operation based on the center point in the expanded rectangle to obtain the target training samples; wherein, the target region is the largest square region in the expanded rectangle, and the center point of the target region coincides with the center point of the expanded rectangle. In the above way, target training samples with dimensions meeting the requirements can be obtained, and the dimensions of the target training samples are set to be square, that is, the aspect ratio is the same, which is more convenient for the model to process the sample data and can effectively avoid problems such as distortion.

[0040] The finally obtained target training samples in the above way are clearer, the proportion of the salient object is larger, and the dimensions of the target training samples are also more convenient for the model to process, further improving the two-stage training effect.

[0041] In summary, for ease of understanding, the embodiments of the present disclosure provide a structural schematic diagram of a custom generation model as shown in Figure 3 which includes a VAE encoding network, a denoising network, a VAE decoding network, a text encoder, and an image encoder. The VAE encoding network, the text encoder, and the image encoder each correspond to input data (image or text). After the output of the VAE encoding network is added with noise, the data Zt with added noise is obtained, and Zt can be denoised through the denoising network. In Figure 3Among them, the denoising network includes multiple denoising units to achieve layer-by-layer denoising. Zt-1 is the output result of the first denoising unit, and so on until Z0 output by the last denoising unit is obtained, so that the VAE decoding network decodes based on Z0 and finally outputs the required image. In addition, it should be noted that the text encoder and the image encoder can respectively extract features from their input data and input the extracted features into the denoising units of the denoising network to achieve information fusion, and finally the VAE decoding network outputs the generated image. During the first-stage training, the parameters of the VAE encoding network, the denoising network, the VAE decoding network, and the text encoder can be fixed, and the parameters of the image encoder and the network parameters involved in the foregoing Adapter unit can be adjusted. During the second-stage training, the parameters of the VAE encoding network, the VAE decoding network, and the image encoder can be fixed, and the parameters of the denoising network, the text encoder, and the network parameters involved in the LoRA technique can be adjusted based on the LoRA technique.

[0042] Through the above method, the training difficulty can be effectively reduced, and the feature extraction ability of the generation model can be better improved, effectively ensuring the feature presentation effect of the trained generation model.

[0043] Corresponding to the foregoing Figure 4 As shown in the structural schematic diagram of a construction device of a generation model provided by an embodiment of the present disclosure, the device can be implemented by software and / or hardware, and is generally integrated in an electronic device, such as Figure 4 As shown, the construction device of the generation model includes:

[0044] A network acquisition module 402, configured to acquire a preset general generation network;

[0045] A model construction module 404, configured to construct an initial custom generation model based on the general generation network and a custom conditional injection network;

[0046] A first-stage training module 406, configured to perform first-stage training on the initial custom generation model to obtain a custom generation model after the first-stage training; wherein, the parameters of the first feature extraction unit in the custom conditional injection network change during the first-stage training;

[0047] A second-stage training module 408, configured to perform second-stage training on the custom generation model after the first-stage training to obtain a custom generation model after the second-stage training; wherein, some network parameters of the general generation network and the parameters of the second feature extraction unit in the custom conditional injection network change during the second-stage training, the parameters of the first feature extraction unit remain unchanged during the second-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0048] The above technical solution provided by the embodiments of the present disclosure can efficiently construct an initial custom generation model on the basis of a preset general generation network by combining a custom conditional injection network. By adopting a two-stage training method for the constructed generation model, the model parameters in different training stages are different, which can effectively reduce the training difficulty. And because the constructed custom generation model introduces a custom conditional injection network on the basis of the general generation network, the feature extraction ability of the generation model can be better improved according to the requirements, effectively ensuring the feature presentation effect of the trained generation model.

[0049] In some embodiments, the general generation network includes an encoding network, a denoising network, and a decoding network connected in sequence; some network parameters of the general generation network are the parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during the first-stage training and the second-stage training.

[0050] In some embodiments, the first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder.

[0051] In some embodiments, the device further includes a sample processing module, configured to obtain an initial training sample used for the first-stage training of the initial custom generation model; perform a cleaning process on the initial training sample to obtain a target training sample; the target training sample is used for the second-stage training of the custom generation model after the first-stage training.

[0052] In some embodiments, the sample processing module is specifically configured to: obtain target information corresponding to the initial training sample; the target information includes the resolution of the salient object in the initial training sample and / or the clarity of the initial training sample; based on the target information, perform a filtering operation on the initial training sample; based on the initial training samples remaining after the filtering operation, obtain a target training sample.

[0053] In some embodiments, when the target information includes the resolution of the salient object in the initial training sample, the sample processing module is specifically configured to: filter the initial training samples whose resolution of the salient object is less than a preset resolution threshold.

[0054] In some embodiments, when the target information includes the clarity of the initial training sample, the sample processing module is specifically configured to: determine a clarity threshold based on the clarity of each initial training sample; filter the initial training samples whose clarity is less than the clarity threshold.

[0055] In some embodiments, the clarity corresponding to the initial training samples is determined based on the Gaussian covariance values of the initial training samples; specifically, the sample processing module is configured to: perform an averaging process based on the Gaussian covariance values corresponding to the initial training samples respectively to obtain an average value; determine a clarity threshold based on the product of the average value and a preset ratio.

[0056] In some embodiments, the sample processing module is specifically configured to: obtain the bounding rectangle of the salient object in the initial training samples remaining after the filtering operation; perform an expanding process on the bounding rectangle to obtain an expanded rectangle; crop a target region from the initial training samples remaining after the filtering operation based on the center point in the expanded rectangle to obtain target training samples; wherein, the target region is the largest square region in the expanded rectangle, and the center point of the target region coincides with the center point of the expanded rectangle.

[0057] The construction device of the generation model provided by the embodiments of the present disclosure can execute the construction method of the generation model provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0058] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device embodiments described above can refer to the corresponding process in the method embodiments, and will not be described herein again.

[0059] The embodiments of the present disclosure provide an electronic device, which includes: a storage device storing a computer program thereon; a processing device configured to execute the computer program in the storage device to implement the steps of any method in the present disclosure.

[0060] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0061] As Figure 5As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0062] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0063] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0064] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to execute the image processing method provided by the embodiments of the present disclosure. The computer program products may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0065] In addition, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, cause the processor to execute the method for constructing a generation model provided by the embodiments of the present disclosure.

[0066] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0067] Embodiments of the present disclosure also provide a computer program product, including computer programs / instructions that, when executed by a processor, implement the method for constructing a generation model in the embodiments of the present disclosure.

[0068] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0069] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.

[0070] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0071] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0072] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0073] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a generative model, characterized in that, Including: Obtain a preset general generation network; Based on the general generation network and a custom conditional injection network, construct an initial custom generation model; Perform a first-stage training on the initial custom generation model to obtain a custom generation model after the first-stage training; wherein, the parameters of the first feature extraction unit in the custom conditional injection network change during the first-stage training; Perform a second-stage training on the custom generation model after the first-stage training to obtain a custom generation model after the second-stage training; wherein, some network parameters of the general generation network and the parameters of the second feature extraction unit in the custom conditional injection network change during the second-stage training, the parameters of the first feature extraction unit remain unchanged during the second-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

2. The method according to claim 1, characterized in that, The general generation network includes an encoding network, a denoising network, and a decoding network connected in sequence; Some network parameters of the general generation network include the parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during both the first-stage training and the second-stage training.

3. The method according to claim 1, wherein The first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder.

4. The method according to claim 1, characterized in that The performing a first-stage training on the initial custom generation model includes: using a preset initial training sample to perform a first-stage training on the initial custom generation model; The performing a second-stage training on the custom generation model after the first-stage training includes: cleaning the initial training sample to obtain a target training sample; using the target training sample to perform a second-stage training on the custom generation model after the first-stage training.

5. The method according to claim 4, characterized in that The cleaning the initial training sample to obtain a target training sample includes: Obtain target information corresponding to the initial training sample; the target information includes the resolution of the significant object in the initial training sample and / or the clarity of the initial training sample; Based on the target information, perform a filtering operation on the initial training sample; Based on the initial training samples remaining after the filtering operation, obtain a target training sample.

6. The method according to claim 5, characterized in that, When the target information includes the resolution of the significant object in the initial training sample, the performing a filtering operation on the initial training sample based on the target information includes: Filtering the initial training samples with the resolution of the significant object less than a preset resolution threshold.

7. The method according to claim 5, wherein When the target information includes the clarity of the initial training sample, the performing a filtering operation on the initial training sample based on the target information includes: Determine a clarity threshold based on the clarity corresponding to each of the initial training samples; Filter the initial training samples with clarity less than the clarity threshold.

8. The method according to claim 7, wherein The clarity corresponding to the initial training sample is determined based on the Gaussian covariance value of the initial training sample; The determining a clarity threshold based on the clarity corresponding to each of the initial training samples includes: Perform an averaging process based on the Gaussian covariance values corresponding to the respective initial training samples to obtain an average value; Determine a clarity threshold based on the product between the average value and a preset ratio.

9. The method according to claim 5, characterized in that, The obtaining of the target training samples based on the initial training samples remaining after the filtering operation includes: Obtain the bounding rectangle of the significant object in the initial training samples remaining after the filtering operation; Perform an expanding process on the bounding rectangle to obtain an expanded rectangle; Based on the center point in the expanded rectangle, crop a target region from the initial training samples remaining after the filtering operation to obtain the target training samples; wherein, the target region is the largest square region in the expanded rectangle, and the center point of the target region coincides with the center point of the expanded rectangle.

10. A construction device for a generation model, characterized in that, Includes: A network acquisition module, configured to acquire a preset general generation network; A model construction module, configured to construct an initial custom generation model based on the general generation network and a custom conditional injection network; A first-stage training module, configured to perform first-stage training on the initial custom generation model to obtain a custom generation model after first-stage training; wherein, the parameters of the first feature extraction unit in the custom conditional injection network change during the first-stage training; A second-stage training module, configured to perform second-stage training on the custom generation model after first-stage training to obtain a custom generation model after second-stage training; wherein, some network parameters of the general generation network and the parameters of the second feature extraction unit in the custom conditional injection network change during the second-stage training, the parameters of the first feature extraction unit remain unchanged during the second-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

11. An electronic device, characterized in that, The electronic device includes: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method for constructing a generation model according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the method for constructing a generation model according to any one of the above claims 1-9.

13. A computer program product, characterized in that, Includes a computer program, and the computer program, when executed by a processor, implements the method for constructing a generation model according to any one of claims 1-9.