Method and apparatus for constructing generative model, device, and medium

By introducing custom conditional injection network and two-stage training methods into the generative model, the shortcomings of the existing generative model in feature presentation effect are solved, and more efficient feature extraction and presentation effect are achieved.

WO2025130777A1PCT designated stage expired Publication Date: 2025-06-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139156
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing generative models are not good in feature presentation, especially in portrait generation tasks, which cannot accurately restore features such as facial facial features and skin types.

Method used

Based on the preset general generation network, the initial custom generation model is constructed by injecting the network with customized conditions, and the model parameters are adjusted through two-stage training to ensure that parameter changes at different stages help to improve feature extraction capabilities.

Benefits of technology

By introducing custom conditional injection networks and two-stage training methods, the feature extraction capability and feature presentation effect of the generative model are significantly improved, ensuring that the generated images are more realistic and diverse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139156_26062025_PF_FP_ABST
    Figure CN2024139156_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and apparatus for constructing a generative model, a device, and a medium. The method comprises: acquiring a preset general-purpose generative network; on the basis of the general-purpose generative network and a customized conditional injection network, constructing an initial customized generative model; performing first-stage training on the initial customized generative model to obtain a customized generative model that has undergone first-stage training, parameters of a first feature extraction unit in the customized conditional injection network being changed during the first-stage training; and performing second-stage training on the customized generative model that has undergone first-stage training to obtain a customized generative model that has undergone second-stage training, some network parameters of the general-purpose generative network and parameters of a second feature extraction unit in the customized conditional injection network being changed during the second-stage training, and the parameters of the first feature extraction unit not being changed during the second-stage training. According to the embodiments of the present disclosure, it is possible to effectively reduce the difficulty of training and ensure the feature presentation effect of the generative model obtained by training.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and medium for constructing generative model

[0001] This application claims priority to the Chinese invention patent application entitled “Method, device, equipment and medium for constructing a generative model” and application number 202311786637.3, filed on December 22, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of computer technology, and in particular to a method, device, equipment, and medium for constructing a generation model. Background Art

[0003] With the increasing popularity of generative models, they are being used in a variety of fields, such as gaming and academia, to quickly and easily generate the required multimedia information. For example, image generative models can be used to generate realistic or diverse images. However, existing generative models have poor feature representation performance. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, device, equipment and medium for constructing a generation model.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for constructing a generative model, the method comprising: obtaining a preset universal generative network; constructing an initial custom generative model based on the universal generative network and a custom conditional injection network; performing one-stage training on the initial custom generative model to obtain a custom generative model after one-stage training; wherein the parameters of the first feature extraction unit in the custom conditional injection network change during the one-stage training; performing two-stage training on the custom generative model after the one-stage training to obtain a custom generative model after the two-stage training; wherein part of the network parameters of the universal generative network and the parameters of the second feature extraction unit in the custom conditional injection network change during the two-stage training, the parameters of the first feature extraction unit remain unchanged during the two-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0006] An embodiment of the present disclosure also provides a device for constructing a generative model, comprising: a network acquisition module for acquiring a preset universal generative network; a model construction module for constructing an initial custom generative model based on the universal generative network and a custom conditional injection network; a one-stage training module for performing one-stage training on the initial custom generative model to obtain a custom generative model after the one-stage training; wherein the parameters of the first feature extraction unit in the custom conditional injection network change during the one-stage training; a two-stage training module for performing two-stage training on the custom generative model after the one-stage training to obtain a custom generative model after the two-stage training; wherein some network parameters of the universal generative network and the parameters of the second feature extraction unit in the custom conditional injection network change during the two-stage training, the parameters of the first feature extraction unit remain unchanged during the two-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0007] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the method for constructing a generative model as provided in an embodiment of the present disclosure.

[0008] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the method for constructing a generation model as provided in the embodiment of the present disclosure.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0011] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] FIG1 is a flow chart of a method for constructing a generative model provided by an embodiment of the present disclosure;

[0013] FIG2 is a schematic diagram of the structure of a custom generation model provided by an embodiment of the present disclosure;

[0014] FIG3 is a schematic diagram of the structure of a custom generation model provided by an embodiment of the present disclosure;

[0015] FIG4 is a schematic structural diagram of a device for constructing a generative model provided by an embodiment of the present disclosure;

[0016] FIG5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The above-mentioned technical solution provided by the embodiment of the present disclosure can efficiently construct an initial customized generative model based on a preset general generative network in combination with a customized conditional injection network. By adopting a two-stage training method for the constructed generative model, the model parameters of different training stages are different, which can effectively reduce the training difficulty. Moreover, since the constructed customized generative model introduces a customized conditional injection network on the basis of the general generative network, the feature extraction capability of the generative model can be better improved according to needs, and the feature presentation effect of the trained generative model can be effectively guaranteed.

[0018] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0019] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0020] Figure 1 is a flow chart of a method for constructing a generative model provided by an embodiment of the present disclosure. The method can be executed by a device for constructing a generative model, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. As shown in Figure 1, the method mainly includes the following steps S102 to S108:

[0021] Step S102: Obtain a preset universal generation network.

[0022] The general generative network can be an already trained generative network, such as an open source generative network, which includes the network structure required for the generative network. For example, the general generative network can adopt an existing general diffusion model. In some implementation examples, the general generative network includes an encoding network, a denoising network, and a decoding network connected in sequence. The embodiments of the present disclosure do not limit the encoding network and the decoding network. For example, the encoding network and the decoding network can be implemented using the encoder and decoder in the VAE (Variational Auto-Encoder) model. The denoising network includes but is not limited to the Unet network. For example, in the image-to-image scenario, the general generative network can be processed based on the input image to generate a stylized image. However, the existing general generative networks are usually unable to restore the detailed features of the input image well. For example, the existing general generative networks cannot accurately restore the facial features, skin texture and other features in the portrait generation task.

[0023] Step S104, based on the universal generation network and the customized conditional injection network, an initial customized generation model is constructed. The disclosed embodiment does not limit the structure of the conditional injection network. For example, it can include functional units such as a feature extraction unit. The conditional injection network can not only extract features from the received data, but also perform processing such as feature mapping (this is only an example and should not be regarded as a limitation. The processing operations to be performed can be flexibly set according to the needs), and the acquired feature-related information and other conditional information are injected into the generation process. Taking the customized generation model to be ultimately constructed as an image generation model as an example, the conditional injection network in the customized generation model can guide image generation to a certain extent through processing operations such as feature extraction and information injection. The conditional injection network is connected to the universal generation network, and different inputs of the model can be correspondingly input to the conditional injection network and the universal generation network. The output result of the conditional injection network can be fused with the intermediate information generated by the universal generation network, and the final generation result of the universal generation network is used as the output result of the customized generation model. For ease of understanding, a structural schematic diagram of a customized generation model shown in Figure 2 can be referred to. In Figure 2, it is assumed that the input of the general generative network is the first input data, and the input of the customized conditional injection network is the second input data. The conditional injection network can perform feature extraction processing on the second input data and fuse the extracted features with the intermediate data generated by the general generative network for the first input data, and ultimately the general generative network outputs the generated data. The embodiments of the present disclosure do not limit the types of input data and generated data. For example, the first input data can be an image, the second input data can include text and images, and the generated data can be an image.

[0024] Step S106: Perform a one-stage training on the initial customized generative model to obtain a customized generative model after the one-stage training. The parameters of the first feature extraction unit in the customized conditional injection network are changed during the one-stage training. For example, the first feature extraction unit may include an image encoder (Img-encoder). In some implementation examples, the parameters of the general generative network remain unchanged during the one-stage training. Furthermore, some parameters of the general generative network may also be changed during the one-stage training as needed, without limitation herein.

[0025] It should also be noted that, in practical applications, the custom generative model may also include an adapter unit. The adapter unit can be implemented using a network such as an MLP (Multilayer Perceptron). For details, please refer to relevant technologies and are not limited here. The adapter unit can map the features of the input data (such as images and / or text) of the conditional injection network to a corresponding feature space. The mapped features can then be injected into the corresponding generation process (such as the image generation process) of the custom generative model via a partial network in the custom generative model. Such partial networks can include the cross-attention unit of a denoising network (e.g., UNet) in a general generative network and the Lora layer added to the custom generative model. In this way, the conditional injection network can inject conditional information, such as feature-related information, into the generation process via the adapter unit, so that the final generated data (such as the generated image) can reflect at least some of the features of the input data to a certain extent. During the training phase, the relevant network parameters of the adapter unit are also adjusted to ensure that the expected information processing capabilities are achieved.

[0026] Step S108: Perform a second-stage training on the customized generative model after the first-stage training to obtain a customized generative model after the second-stage training; wherein, during the second-stage training, some network parameters of the general generative network and the parameters of the second feature extraction unit in the customized conditional injection network are changed, while the parameters of the first feature extraction unit remain unchanged during the second-stage training; the first feature extraction unit and the second feature extraction unit are used to extract different feature types, including but not limited to text features and image features. In some specific implementation examples, the first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder.

[0027] In some specific implementation examples, some network parameters of the universal generative network include parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during both the first-stage training and the second-stage training, and the parameters of the first feature extraction unit remain unchanged during the second-stage training. For example, when the denoising network is a Unet network, the parameters of the Unet network are only changed during the second-stage training. Furthermore, the second-stage training is also used to adjust the parameters of the second feature extraction unit. For example, the parameters of the text encoder are changed in the second stage, but the parameters of the image encoder have already been adjusted in the first stage. Therefore, during the second-stage training, the parameters of the image encoder remain unchanged.

[0028] In order to ensure the effect of the two-stage training, in the embodiment of the present disclosure, the custom generation model after the one-stage training can be trained in the second stage based on LoRA technology (also called LoRA fine-tuning technology or LoRA training technology). For example, the parameters of the denoising network (such as the Unet network) and the text encoder are adjusted based on LoRA technology. LoRA technology can assist in model training by inserting some network layers (which can be called LoRA layers) into the model to be trained. During the two-stage training, the network parameters involved in the LoRA technology will also be adjusted accordingly. For details, please refer to the relevant technology and will not be repeated here.

[0029] The above-mentioned technical solution provided by the embodiment of the present disclosure can efficiently construct an initial customized generative model based on a preset general generative network in combination with a customized conditional injection network. By adopting a two-stage training method for the constructed generative model, the model parameters of different training stages are different, which can effectively reduce the training difficulty. Moreover, since the constructed customized generative model introduces a customized conditional injection network on the basis of the general generative network, the feature extraction capability of the generative model can be better improved according to needs, and the feature presentation effect of the trained generative model can be effectively guaranteed.

[0030] In some implementation examples, a first-stage training of an initial custom generative model includes: using a preset initial training sample to perform the first-stage training on the initial custom generative model; based on this, a second-stage training of the custom generative model after the first-stage training includes: cleaning the initial training sample to obtain a target training sample; and using the target training sample to perform the second-stage training on the custom generative model after the first-stage training. Compared to related art methods that directly continue the second-stage training with the training sample used in the first-stage training, the embodiments of the present disclosure can clean the initial training sample to obtain high-quality training samples that are more suitable for the second-stage training, thereby improving the effect of the second-stage training.

[0031] The present disclosure provides an implementation example of cleaning the initial training samples to obtain target training samples, which can be performed with reference to the following steps A to C:

[0032] Step A: Obtain target information corresponding to the initial training sample; the target information includes the resolution of salient objects in the initial training sample and / or the clarity of the initial training sample. In practical applications, if the target information includes the resolution of salient objects, a salient object detection algorithm (or salient detection model) can be first used to detect salient objects in the initial training sample. Salient objects can be, for example, people in an image. The disclosed embodiments do not limit the type of salient objects.

[0033] In step B, the initial training samples are filtered based on the target information. Specifically, the initial training samples with poor target information can be filtered out, thereby retaining high-quality samples.

[0034] For example, if the target information includes the resolution of salient objects in the initial training samples, initial training samples with salient object resolutions below a preset resolution threshold can be filtered out. For example, the preset resolution threshold can be 100*100 pixels. This effectively ensures that the remaining sample images have a desired proportion of salient objects.

[0035] In the case where the target information includes the clarity of the initial training samples, a clarity threshold may be determined based on the clarity corresponding to each of the initial training samples; and then initial training samples with clarity less than the clarity threshold may be filtered.

[0036] In some specific implementation examples, the clarity of the initial training samples is determined based on the Gaussian covariance values ​​of the initial training samples. Determining a clarity threshold based on the clarity of each of the initial training samples includes: averaging the Gaussian covariance values ​​of the initial training samples to obtain an average value; and determining the clarity threshold based on the product of the average value and a preset ratio. The preset ratio can be flexibly set as needed, such as to 10%. This approach effectively ensures that the remaining sample images have a high degree of clarity.

[0037] Step C: Obtain a target training sample based on the initial training sample remaining after the filtering operation. In some implementation examples, the initial training sample remaining after the filtering operation can be directly used as the target training sample. In other implementation examples, further operations can be performed on the initial training sample remaining after the filtering operation, such as cropping a key area therefrom, to obtain a target training sample that meets the requirements. For example, steps C1 to C3 can be performed as follows:

[0038] Step C1: Obtain the bounding rectangle of the salient object in the initial training sample remaining after the filtering operation. The bounding rectangle of the salient object can be obtained based on a salient object detection algorithm.

[0039] Step C2: Expand the circumscribed rectangular frame to obtain an expanded rectangular frame. In practical applications, the circumscribed rectangular frame may be expanded until its area or size is expanded N times (N is a positive integer, such as 3), or until it reaches the boundary of the initial training sample.

[0040] In step C3, based on the center point of the expanded rectangular frame, a target region is cropped from the initial training samples remaining after the filtering operation to obtain a target training sample. The target region is the largest square region within the expanded rectangular frame, and the center point of the target region coincides with the center point of the expanded rectangular frame. This method can obtain target training samples of the required size. The target training samples are set to a square size with a consistent aspect ratio, which makes it easier for the model to process the sample data and effectively avoids problems such as distortion.

[0041] The target training samples obtained through the above method are clearer, with a larger proportion of salient objects. The size of the target training samples is also easier for the model to process, further improving the effect of the second-stage training.

[0042] In summary, for ease of understanding, the embodiment of the present disclosure provides a structural diagram of a custom generation model as shown in Figure 3, including a VAE encoding network, a denoising network, a VAE decoding network, a text encoder and an image encoder. The VAE encoding network, the text encoder and the image encoder each correspond to input data (image or text). After the output of the VAE encoding network is added with noise, the noisy data Zt is obtained, and Zt can be denoised by the denoising network. In Figure 3, the denoising network includes multiple denoising units, thereby achieving layer-by-layer denoising. Zt-1 is the output result of the first denoising unit, and so on, until Z0 output by the last denoising unit is obtained, so that the VAE decoding network is decoded based on Z0, and finally outputs the desired image. It should also be noted that the text encoder and the image encoder can respectively extract features from their input data, and input the extracted features into the denoising unit of the denoising network, thereby achieving information fusion, and finally the generated image is output by the VAE decoding network. During the first-stage training, the parameters of the VAE encoding network, denoising network, VAE decoding network and text encoder can be fixed, and the parameters of the image encoder and the network parameters involved in the aforementioned Adapter unit can be adjusted. During the second-stage training, the parameters of the VAE encoding network, VAE decoding network and image encoder can be fixed, and the parameters of the denoising network, text encoder and the network parameters involved in the LoRA technology can be adjusted based on the LoRA technology.

[0043] Through the above method, the training difficulty can be effectively reduced, and the feature extraction ability of the generative model can be greatly improved, effectively ensuring the feature presentation effect of the generative model obtained through training.

[0044] Corresponding to FIG4 , which is a schematic diagram of the structure of a device for constructing a generative model provided in an embodiment of the present disclosure, the device may be implemented by software and / or hardware and may generally be integrated into an electronic device. As shown in FIG4 , the device for constructing a generative model includes:

[0045] The network acquisition module 402 is used to acquire a preset universal generation network;

[0046] A model building module 404 is used to build an initial customized generation model based on the general generation network and the customized conditional injection network;

[0047] A one-stage training module 406 is configured to perform one-stage training on the initial customized generative model to obtain a customized generative model after the one-stage training; wherein the parameters of the first feature extraction unit in the customized conditional injection network are changed during the one-stage training;

[0048] The two-stage training module 408 is used to perform two-stage training on the customized generative model after the first-stage training to obtain the customized generative model after the two-stage training; wherein, some network parameters of the general generative network and the parameters of the second feature extraction unit in the customized conditional injection network are changed during the two-stage training, the parameters of the first feature extraction unit remain unchanged during the two-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

[0049] The above-mentioned technical solution provided by the embodiment of the present disclosure can efficiently construct an initial customized generative model based on a preset general generative network in combination with a customized conditional injection network. By adopting a two-stage training method for the constructed generative model, the model parameters of different training stages are different, which can effectively reduce the training difficulty. Moreover, since the constructed customized generative model introduces a customized conditional injection network on the basis of the general generative network, the feature extraction capability of the generative model can be better improved according to needs, and the feature presentation effect of the trained generative model can be effectively guaranteed.

[0050] In some embodiments, the universal generation network includes an encoding network, a denoising network, and a decoding network connected in sequence; some network parameters of the universal generation network are parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during the one-stage training and the two-stage training.

[0051] In some embodiments, the first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder.

[0052] In some embodiments, the device also includes a sample processing module for obtaining initial training samples used for the first-stage training of the initial custom generation model; cleaning the initial training samples to obtain target training samples; and the target training samples are used to perform second-stage training on the custom generation model after the first-stage training.

[0053] In some embodiments, the sample processing module is specifically used to: obtain target information corresponding to the initial training sample; the target information includes the resolution of the salient object in the initial training sample and / or the clarity of the initial training sample; based on the target information, perform a filtering operation on the initial training sample; and obtain a target training sample based on the initial training sample remaining after the filtering operation.

[0054] In some embodiments, when the target information includes the resolution of the salient objects in the initial training samples, the sample processing module is specifically configured to filter the initial training samples in which the resolution of the salient objects is less than a preset resolution threshold.

[0055] In some embodiments, when the target information includes the clarity of the initial training samples, the sample processing module is specifically used to: determine a clarity threshold based on the clarity corresponding to each of the initial training samples; and filter initial training samples whose clarity is less than the clarity threshold.

[0056] In some embodiments, the clarity corresponding to the initial training sample is determined based on the Gaussian covariance value of the initial training sample; the sample processing module is specifically used to: perform averaging processing based on the Gaussian covariance values ​​corresponding to each of the initial training samples to obtain an average value; and determine a clarity threshold based on the product between the average value and a preset ratio.

[0057] In some embodiments, the sample processing module is specifically used to: obtain an enclosed rectangular frame of a significant object in the initial training sample remaining after the filtering operation; perform an expansion process on the enclosed rectangular frame to obtain an enlarged rectangular frame; based on the center point in the enlarged rectangular frame, crop a target area from the initial training sample remaining after the filtering operation to obtain a target training sample; wherein, the target area is the largest square area in the enlarged rectangular frame, and the center point of the target area coincides with the center point of the enlarged rectangular frame.

[0058] The device for constructing a generative model provided in the embodiments of the present disclosure can execute the method for constructing a generative model provided in any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.

[0059] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device embodiment can refer to the corresponding process in the method embodiment, and will not be repeated here.

[0060] An embodiment of the present disclosure provides an electronic device, which includes: a storage device storing a computer program; and a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.

[0061] Reference is now made to FIG5 , which illustrates a schematic diagram of the structure of an electronic device 500 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG5 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.

[0062] As shown in Figure 5, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0063] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0064] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0065] In addition to the above-mentioned methods and devices, the embodiments of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the image processing method provided by the embodiments of the present disclosure. The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present disclosure, the programming languages ​​including object-oriented programming languages ​​such as Java, C++, etc., and also conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0066] In addition, the embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the method for constructing a generation model provided by the embodiment of the present disclosure.

[0067] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0068] The embodiments of the present disclosure also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the method for constructing a generative model in the embodiments of the present disclosure.

[0069] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0070] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0071] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0072] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.

[0074] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a generative model, comprising: Get the preset general generation network; Build an initial custom generation model based on the general generation network and the custom conditional injection network; Performing one-stage training on the initial custom generative model to obtain a one-stage trained custom generative model; wherein the parameters of the first feature extraction unit in the custom conditional injection network are changed during the one-stage training; and The customized generative model after the one-stage training is subjected to a two-stage training to obtain a customized generative model after the two-stage training; wherein, some network parameters of the universal generative network and the parameters of the second feature extraction unit in the customized conditional injection network are changed during the two-stage training, the parameters of the first feature extraction unit remain unchanged during the two-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

2. The method according to claim 1, wherein: The universal generation network includes an encoding network, a denoising network and a decoding network connected in sequence; Part of the network parameters of the universal generation network include parameters of the denoising network, and the parameters of the encoding network and the decoding network remain unchanged during the one-stage training and the two-stage training.

3. The method according to claim 1, wherein: The first feature extraction unit includes an image encoder, and the second feature extraction unit includes a text encoder.

4. The method according to claim 1, wherein: The performing one-stage training on the initial custom generation model comprises: using a preset initial training sample to perform one-stage training on the initial custom generation model; The second-stage training of the custom generated model after the first-stage training includes: cleaning the initial training samples to obtain target training samples; and using the target training samples to perform second-stage training on the custom generated model after the first-stage training.

5. The method according to claim 4, wherein: The cleaning process of the initial training sample to obtain a target training sample includes: Acquiring target information corresponding to the initial training sample; the target information includes the resolution of the salient object in the initial training sample and / or the clarity of the initial training sample; Based on the target information, filtering the initial training samples; Based on the initial training samples remaining after the filtering operation, target training samples are obtained.

6. The method according to claim 5, wherein: In a case where the target information includes the resolution of the salient object in the initial training sample, the filtering operation on the initial training sample based on the target information includes: Initial training samples whose resolution of salient objects is less than a preset resolution threshold are filtered.

7. The method according to claim 5, wherein: In a case where the target information includes the clarity of the initial training sample, the filtering operation on the initial training sample based on the target information includes: Determining a clarity threshold based on the clarity corresponding to each of the initial training samples; Initial training samples whose clarity is less than the clarity threshold are filtered.

8. The method according to claim 7, wherein: The clarity corresponding to the initial training sample is determined based on the Gaussian covariance value of the initial training sample; The determining of the clarity threshold based on the clarity corresponding to each of the initial training samples includes: Performing averaging processing based on the Gaussian covariance values ​​corresponding to the initial training samples to obtain an average value; A clarity threshold is determined based on the product of the average value and a preset ratio.

9. The method according to claim 5, wherein: The step of obtaining a target training sample based on the initial training sample remaining after the filtering operation comprises: Obtaining a bounding rectangular frame of a salient object in the initial training sample remaining after the filtering operation; Performing an outward expansion process on the circumscribed rectangular frame to obtain an outward expanded rectangular frame; Based on the center point in the outward-expanding rectangular frame, a target area is cropped from the initial training samples remaining after the filtering operation to obtain a target training sample; wherein the target area is the largest square area in the outward-expanding rectangular frame, and the center point of the target area coincides with the center point of the outward-expanding rectangular frame.

10. A device for constructing a generative model, comprising: A network acquisition module, used to acquire a preset universal generation network; A model building module, used to build an initial customized generative model based on a general generative network and a customized conditional injection network; A one-stage training module, used for performing one-stage training on the initial custom generation model to obtain a custom generation model after one-stage training; wherein the parameters of the first feature extraction unit in the custom conditional injection network are changed during the one-stage training; A two-stage training module is used to perform two-stage training on the custom generative model after the one-stage training to obtain the custom generative model after the two-stage training; wherein, some network parameters of the universal generative network and the parameters of the second feature extraction unit in the custom conditional injection network are changed during the two-stage training, the parameters of the first feature extraction unit remain unchanged during the two-stage training, and the first feature extraction unit and the second feature extraction unit are used to extract different types of features.

11. An electronic device, comprising: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method for constructing a generation model as described in any one of claims 1-9.

12. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method for constructing a generation model as described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for constructing a generation model according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Face mask editing method based on generative adversarial network

    CN115439311A

  • Image generation method based on progressive growth condition generative adversarial network

    CN115439323A

  • Data processing method and device

    CN117173035A

  • Generative network based probabilistic portfolio management

    US20210027379A1

  • Autonomous vehicle perception multimodal sensor data management

    US20220114805A1