A Method for Constructing an Intelligent Image Generation Model for Diversified Industry Applications

Through the diversified training and adversarial training methods of intelligent image generation models, the problem of single use scenarios and complex training in the existing technology is solved, and the widespread use and image quality improvement in diversified industry application scenarios is achieved.

CN119338942BActive Publication Date: 2025-06-20XIAMEN SHIBAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411850077.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-06-20
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

In the prior art, the use scenarios of the image generation model are relatively single, and it is difficult to adapt to diversified industry application scenarios, and the training process is complex, making it difficult to adapt to other usage scenarios.

Method used

By using intelligent image generation models, trained based on different industry categories and demand description information, images that meet industry applications are generated, and loss functions are determined using adversarial training methods to improve the authenticity and rendering effect of the model.

Benefits of technology

It realizes the widespread use of intelligent image generation models in diversified industry application scenarios, can quickly build models that adapt to usage scenarios in various industries, and improves the authenticity and rendering effect of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119338942B_ABST
    Figure CN119338942B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing an intelligent image generation model for diversified industry applications, which relates to the technical field of computer models. The method includes: obtaining a first description feature vector according to the first industry category information and the first requirement description information, and inputting it into the intelligent image generation model to obtain a first training generated image; performing adversarial training on the first training generated image, the first sample image set and the discriminant model group to obtain a trained intelligent image generation model; inputting the second industry category information and the second requirement description information into the trained intelligent image generation model to obtain a second training generated image, and inputting the second requirement description information into a dedicated image generation model to obtain a third training generated image, and training the dedicated image generation model. According to the present invention, the intelligent image generation model can generate images that meet industry applications, and a lightweight dedicated image generation model suitable for the usage scenarios of various industries can be quickly constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer models, and particularly to a method for constructing an intelligent image generation model for diversified industrial applications. Background Art

[0002] In the related art, CN118839789A discloses a model training method, an image generation method, a device and an electronic device. In this method, first, a base image and a first description text are obtained, and the base image and the first description text are input into an image generation model to be trained, so that the image generation model determines the image features corresponding to the base image and the text features corresponding to the first description text, and based on the image features corresponding to the base image and the text features corresponding to the first description text, generates an image of a target object with the physical characteristics of a reference object in a specified environment as an output image, determines a comprehensive loss function value according to the feature deviation between the image features corresponding to the output image and the image features corresponding to the base image, and the similarity between the feature of the image content expressed by the output image and the text features corresponding to the first description text, and trains the image generation model according to the comprehensive loss function value.

[0003] CN118114788A discloses a training method of an image generation model, an image generation method and related devices. The method includes: obtaining a plurality of training data pairs; for the first training data pair, based on a text understanding module, understanding the first training text in the first training data pair to obtain a first vector; interacting the first vector with the embedding vector of the first training text to obtain a second vector; generating a third vector of the input feature dimension required by the image generation module based on the second vector and a dimension mapping module; generating a target image corresponding to the first training text based on the image generation module and the third vector; determining a target loss based on the target image corresponding to each training data pair, the positive correlation image and the negative correlation image in each training data pair; freezing the parameters of the image generation module based on the target loss, and training the text understanding module and the dimension mapping module to obtain an image generation model.

[0004] Therefore, in the related art, although an image generation model can be trained and text information can be used to generate images, the usage scenarios of the model are relatively single, it is difficult to be used in diversified industrial application scenarios, and due to the complexity of the training process, it is difficult for the model to adapt to other usage scenarios.

[0005] The information disclosed in the background art part of the present application is only intended to deepen the understanding of the general background art of the present application, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0006] The present invention provides a method for constructing an intelligent image generation model for diversified industrial applications, which can solve the technical problems in the related art that the usage scenarios of the image generation model are single and it is difficult to adapt to other usage scenarios.

[0007] According to a first aspect of the present invention, there is provided a method for constructing an intelligent image generation model for diversified industrial applications, including:

[0008] Obtain a first description feature vector according to first industry category information and first requirement description information, wherein the first industry category information is used to describe the industry category corresponding to the generated image, and the first requirement description information is used to describe the content included in the generated image;

[0009] Process the first description feature vector by an intelligent image generation model to obtain a first training generated image;

[0010] According to the first training generated image, a first sample image set including first sample images of multiple industry categories, and a discriminant model group, obtain a combined loss function of the intelligent image generation model and the discriminant model group;

[0011] Train the intelligent image generation model through the combined loss function to obtain a trained intelligent image generation model;

[0012] Obtain a second training generated image according to second industry category information, second requirement description information, and the trained intelligent image generation model;

[0013] Input the second requirement description information into a dedicated image generation model corresponding to the second industry category information to obtain a third training generated image;

[0014] Obtain a dedicated loss function of the dedicated image generation model according to the third training generated image and the second training generated image;

[0015] Train the dedicated image generation model according to the dedicated loss function to obtain a trained dedicated image generation model.

[0016] Technical effects: According to the present invention, the intelligent image generation model can be trained by the first industry category information and the first demand description information of multiple industry categories, so that the intelligent image generation model can generate images that meet industry applications, facilitate use in diversified industry application scenarios, and the dedicated image generation models of each industry can be trained by using the intelligent image generation model, facilitating the construction of lightweight dedicated image generation models based on industries and fields, so that models adapted to the usage scenarios of various industries can be quickly constructed. The first authenticity loss function can be determined through adversarial training, so as to be able to improve the authenticity of the images generated by the intelligent image generation model while enhancing the discrimination ability of the authenticity discrimination model, thereby being able to balance the performance of the two models and improve the authenticity of the first training generated images. The first industry loss function can be determined through adversarial training, so as to be able to improve the ability of the intelligent image generation model to describe industry characteristics and the accuracy of the rendering effect while enhancing the ability of the industry category discrimination model to judge the industry category to which the image belongs, thereby being able to balance the performance of the two models and further improve the authenticity of the first training generated images and the accuracy of the rendering effect. The first content loss function can be determined through adversarial training, so as to be able to improve the accuracy of the content generated by the intelligent image generation model while enhancing the ability of the content category discrimination model to judge whether the content of the image is correct, thereby being able to balance the performance of the two models and further improve the authenticity of the first training generated images and the accuracy of the included content, making the content in the images generated by the intelligent image generation model more in line with expectations. The dedicated loss function can be determined based on the error of the pixel values of the third training generated image and the second training generated image at each pixel point, so as to improve the consistency between the images generated by the dedicated image generation model and the images of the industry generated by the intelligent image generation model, enabling the dedicated image generation model to achieve an effect similar to that of the intelligent image generation model when generating images of its corresponding industry and being more lightweight, facilitating construction and training.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present invention. Other features and aspects of the present invention will become clearer according to the following detailed description of the exemplary embodiments with reference to the accompanying drawings. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other embodiments can be obtained based on these drawings without creative efforts.

[0019] Figure 1A flowchart showing a method for constructing an intelligent image generation model for diversified industry applications according to an embodiment of the present invention is exemplarily shown. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 A flowchart showing a method for constructing an intelligent image generation model for diversified industry applications according to an embodiment of the present invention is exemplarily shown. The method includes:

[0023] Step S101: Obtain a first description feature vector according to first industry category information and first requirement description information, where the first industry category information is used to describe the industry category corresponding to the generated image, and the first requirement description information is used to describe the content included in the generated image;

[0024] Step S102: Process the first description feature vector by an intelligent image generation model to obtain a first training generated image;

[0025] Step S103: Obtain a combined loss function of the intelligent image generation model and a discriminant model group according to the first training generated image, a first sample image set including first sample images of multiple industry categories, and the discriminant model group;

[0026] Step S104: Train the intelligent image generation model through the combined loss function to obtain a trained intelligent image generation model;

[0027] Step S105: Obtain a second training generated image according to second industry category information, second requirement description information, and the trained intelligent image generation model;

[0028] Step S106: Input the second requirement description information into a dedicated image generation model corresponding to the second industry category information to obtain a third training generated image;

[0029] Step S107: Obtain the special loss function of the special image generation model based on the third training-generated image and the second training-generated image.

[0030] Step S108: Train the special image generation model according to the special loss function to obtain the trained special image generation model.

[0031] According to the method for constructing an intelligent image generation model for diversified industry applications according to an embodiment of the present invention, the intelligent image generation model can be trained through the first industry category information and the first requirement description information of multiple industry categories, so that the intelligent image generation model can generate images that meet industry applications, facilitating use in diversified industry application scenarios. Moreover, the special image generation models of each industry can be trained using the intelligent image generation model, facilitating the construction of lightweight special image generation models based on industries and fields, thereby enabling the rapid construction of models adapted to the usage scenarios of various industries.

[0032] According to an embodiment of the present invention, in step S101, the first industry category information can be used to represent the industry category to which the generated image belongs. For example, whether the generated image is for advertising media or scientific research, etc. The requirements for the image are different. For an image used in advertising media, it is required to reflect the beauty of the target object in the image. Therefore, parameters such as brightness, contrast, and saturation need to be adjusted to render the image and enhance its beauty. While for an image used in scientific research, the main requirement is to enhance the authenticity of the image, being able to truly reflect the situation of the target object in the experimental scenario. The first industry category information can be text information used to describe the industry category. The first requirement description information is used to describe the content included in the generated image, that is, the target object, the number of target objects, the surrounding environmental background, etc. included in the image. The first requirement description information can also be text information. The first description feature vector can be obtained through a natural language processing model for the first industry category information and the first requirement description information. The natural language processing model can be a deep learning neural network model such as a recurrent neural network. The present invention does not limit the specific type of the natural language processing model.

[0033] According to an embodiment of the present invention, in step S102, the intelligent image generation model is a deep learning neural network model such as a convolutional neural network model, which can generate images based on the input vector. Inputting the above first description feature vector for describing the industry category and the included content into the intelligent image generation model can generate the first training-generated image. There may be differences between the content included in the first training-generated image and the applicable industry and the first requirement description information and the first industry category information. Therefore, the intelligent image generation model can be trained to narrow the above differences.

[0034] According to an embodiment of the present invention, in step S103, the discrimination model group may include multiple discrimination models, which can be trained adversarially with the intelligent image generation model. That is, the discrimination model group can improve the discrimination accuracy during training, and the intelligent image generation model can improve the authenticity and accuracy of the generated images during training.

[0035] According to an embodiment of the present invention, the discrimination model group includes a authenticity discrimination model, an industry category discrimination model, and a content category discrimination model. Among them, the authenticity discrimination model is used to judge the authenticity of the input image, the industry category discrimination model is used to judge the industry category corresponding to the input image, and the content category discrimination model is used to judge the category of the content included in the input image.

[0036] According to an embodiment of the present invention, based on the first training generated images, a first sample image set including first sample images of multiple industry categories, and the discrimination model group, a combined loss function of the intelligent image generation model and the discrimination model group is obtained, including: inputting the first training generated images or the first sample images into the authenticity discrimination model to obtain a first authenticity discrimination result; determining a first authenticity loss function of the authenticity discrimination model according to the first authenticity discrimination result; inputting the first training generated images or the first sample images into the industry category discrimination model to obtain a first industry category description vector; obtaining a first industry loss function of the industry category discrimination model according to the first industry category description vector, the first industry category information, and the industry category of the first sample image; inputting the first training generated images or the first sample images into the content category discrimination model to obtain a first content description vector; obtaining a first content loss function of the content category discrimination model according to the first content description vector, the first demand description information, and the content annotation information of the first sample image; and obtaining the combined loss function according to the first authenticity loss function, the first industry loss function, and the first content loss function.

[0037] According to an embodiment of the present invention, the first sample image is a real image, and the first training generated image is a generated image. Therefore, theoretically, the authenticity discrimination model should determine that the probability of the first sample image being a real image is 100%, and the probability of the first training generated image being a real image is 0. However, as the training process progresses, the realism of the intelligent image generation model becomes higher and higher, resulting in an increase in the probability that the authenticity discrimination model determines the first training generated image to be a real image. Moreover, as the authenticity discrimination model also improves its discrimination ability during training, it can again lead to a decrease in the probability that the authenticity discrimination model determines the first training generated image to be a real image. Eventually, the performance of the authenticity discrimination model and the authenticity discrimination model can reach a balance. Thus, even when the discrimination accuracy of the authenticity discrimination model is very high, it is still difficult to determine whether the first training generated image generated by the intelligent image generation model is a real image, that is, the generated image by the intelligent image generation model has a high degree of realism.

[0038] According to an embodiment of the present invention, determining the first authenticity loss function of the authenticity discrimination model according to the first authenticity discrimination result includes: determining the first authenticity loss function of the authenticity discrimination model according to formula (1) ,

[0039] (1)

[0040] wherein, is the i-th first sample image, is the probability that the authenticity discrimination model outputs that the i-th first sample image is a real image, is the j-th first training generated image, is the probability that the authenticity discrimination model outputs that the j-th first training generated image is a real image, is the number of first sample images in a training batch, is the number of first training generated images in a training batch.

[0041] According to an embodiment of the present invention, according to the above training method that balances the performance of the intelligent image generation model and the authenticity discrimination model, the authenticity discrimination model can be adjusted in the direction of maximizing the first authenticity loss function represented by formula (1), and the intelligent image generation model can be adjusted in the direction of minimizing the first authenticity loss function.

[0042] According to an embodiment of the present invention, in formula (1), if the input image is the first sample image, then during the training of the authenticity discrimination model, make maximize, that is, as close to 1 as possible, so as to enable Maximize it to improve the accuracy of the authenticity discrimination model when judging real images. If the input image is the first training generated image, during the training of the authenticity discrimination model, make Minimize it, that is, make it as close to 0 as possible, so that Can be maximized, thereby improving the accuracy of the authenticity discrimination model when judging generated images.

[0043] According to an embodiment of the present invention, on the other hand, in formula (1), if the input image is the first training generated image, during the training of the intelligent image generation model, make Maximize it, that is, make it as close to 1 as possible, so that Can be minimized, improving the probability that the first training generated image generated by the intelligent image generation model is recognized as a real image by the authenticity discrimination model, thereby improving the realism of the images generated by the intelligent image generation model and enhancing the performance of the intelligent image generation model.

[0044] In this way, through adversarial training, the first authenticity loss function can be determined, so as to meet the purpose of improving the realism of the images generated by the intelligent image generation model while enhancing the discrimination ability of the authenticity discrimination model, thereby balancing the performance of the two models and improving the realism of the first training generated image.

[0045] According to an embodiment of the present invention, the industry category discrimination model can analyze information such as the style, brightness, contrast, and saturation of the first training generated image or the first sample image, so as to judge the industry field to which the image is applied and output the first industry category description vector, and the first industry category description vector can be used to describe the industry category.

[0046] According to an embodiment of the present invention, based on the first industry category description vector, the first industry category information, and the industry category of the first sample image, obtaining the first industry loss function of the industry category discrimination model includes: obtaining the rendering parameter vectors of multiple first sample images of each industry category in the first sample image set, where each component of the rendering parameter vector is a rendering parameter, and the rendering parameter includes the light and shadow contrast parameter, the saturation parameter, the brightness parameter, and the area ratio of the target area to the background area; determining the average rendering parameter vector of each industry category according to the rendering parameter vectors of the multiple first sample images; obtaining the training rendering parameter vector of the first training generated image or the first sample image input into the industry category discrimination model; and determining the first industry loss function of the industry category discrimination model according to the training rendering parameter vector, the average rendering parameter vector, the first industry category description vector, and the first industry category information.

[0047] According to an embodiment of the present invention, the style of the images applied to each industry category can be determined, the rendering parameter vectors of the first sample images of each industry category are obtained, and each component of the rendering parameter vector is a rendering parameter, and the rendering parameter includes a light and shadow contrast parameter, a saturation parameter, a brightness parameter, and the area ratio of the target area to the background area. There are significant differences in the rendering parameters of the images used in different industries. For example, the objects included in the images used in the media industry and scientific research may be the same, but their rendering parameters are different. In order to improve the aesthetics, the images used in the media industry can use parameters such as higher contrast and saturation, while the images used in scientific research need to minimize the difference between the objects in the images and the direct observation of the objects by the human eye and maximize the authenticity. Therefore, the rendering parameter vectors of the first sample images used in different industries may be different from each other. Moreover, the average values of the rendering parameter vectors of multiple first sample images of each industry category can be taken to obtain the average rendering parameter vector of each industry category, so as to reflect the average level of the rendering parameters of the images used in each industry category. Further, the training rendering parameter vector of the first training generated image or the first sample image can be obtained for comparison with the above average rendering parameter vector.

[0048] According to an embodiment of the present invention, based on the training rendering parameter vector, the average rendering parameter vector, the first industry category description vector, and the first industry category information, a first industry loss function of the industry category discrimination model is determined, including: determining the first industry loss function of the industry category discrimination model according to formula (2) ,

[0049] (2)

[0050] wherein, is the first industry category description vector of the i-th first sample image output by the industry category discrimination model, is the transposed vector of, is the category description vector determined according to the industry category of the i-th first sample image, is the first industry category description vector of the j-th first training generated image output by the industry category discrimination model, is the transposed vector of, is the category description vector determined according to the first industry category information of the j-th first training generated image, is the training rendering parameter vector of the j-th first training generated image, is the transposed vector of, is the average rendering parameter vector of the first industry category of the j-th first training generated image.

[0051] According to an embodiment of the present invention, the above-mentioned adversarial training method can still be used. In formula (2), if the input image is the first sample image, then in the training of the industry category discrimination model, Maximize, that is, maximize the similarity between the processing result of the industry category discrimination model on the first sample image and the category description vector of the industry category to which the first sample image belongs, that is, as close to 1 as possible, so that Maximize, thereby improving the accuracy of the industry classification model in determining the industry category to which the real image belongs. If the input image is the first training generated image, in the training of the industry classification model, minimize, that is, as close to 0 as possible, so that Maximize, so that when the industry category discrimination model determines that the generated image belongs to an industry category, it determines that it is a generated image and does not belong to any industry, thereby improving the accuracy of the industry category discrimination model, where, is the similarity between the processing result of the industry category discrimination model on the first training generated image and the category description vector determined based on the first industry category information of the first training generated image, is the probability that the first training generated image is a real image, The similarity between the training rendering parameter vector of the first training generated image and the average rendering parameter vector of the first industry category of the first training generated image is made so that the above product is as close to 0 as possible, that is, even if the authenticity discrimination model determines that the probability that the first training generated image is a real image is not 0, and there is a certain similarity between the training rendering parameter vector and the average rendering parameter vector of the industry to which it belongs, the industry category discrimination model can still determine that the first training generated image is not similar to the first industry category information, that is, output the first industry category description vector that does not belong to any industry, so that it is not similar to the category description vector determined according to the first industry category information.

[0052] According to one embodiment of the present invention, in formula (2), if the input image is the first training generated image, then in the training of the intelligent image generation model, maximize, that is, as close to 1 as possible, so that Minimize, thereby further improving the probability that the first training generated image is a real image and the similarity between the training rendering parameter vector and the average rendering parameter vector. The value of makes the first industry category description vector of the first training generated image closer to the first industry category information, thereby improving the ability of the images generated by the intelligent image generation model to more accurately describe the industry characteristics and obtain more accurate rendering effects, thereby improving the performance of the intelligent image generation model and improving the authenticity of the images generated by the intelligent image generation model.

[0053] In this way, through adversarial training, the first industry loss function can be determined, so that while improving the ability of the intelligent image generation model to describe industry characteristics and the accuracy of the rendering effect, the ability of the industry category discrimination model to judge the industry category to which the image belongs can be improved, thereby balancing the performance of the two models and further improving the authenticity of the first training generated image and the accuracy of the rendering effect.

[0054] According to an embodiment of the present invention, the content category discrimination model can be used to judge whether the content included in the input first training generated image or the first sample image is correct.

[0055] According to an embodiment of the present invention, the first content description vector includes a first target type description vector, a first target state description vector, and a first target scene description vector, and the first demand description information includes first target type description information, first target state description information, and first target scene description information; according to the first content description vector, the first demand description information, and the content annotation information of the first sample image, obtaining the first content loss function of the content category discrimination model includes: obtaining the first content loss function of the content category discrimination model according to formula (3) ,

[0056] (3)

[0057] wherein, is the first target type description vector of the i-th first sample image output by the content category discrimination model, is the transposed vector of, is the target type description vector determined according to the content annotation information of the i-th first sample image, is the first target state description vector of the i-th first sample image output by the content category discrimination model, is the transposed vector of, is the target state description vector determined according to the content annotation information of the i-th first sample image, is the first target scene description vector of the i-th first sample image output by the content category discrimination model, is the transposed vector of, is the target scene description vector determined according to the content annotation information of the i-th first sample image, is the first target type description vector of the j-th first training generated image output by the content category discrimination model, is the transposed vector of, is the target type description vector determined according to the first target type description information of the j-th first training generated image, is the first target state description vector of the j-th first training generated image output by the content category discrimination model, is the transposed vector of, is the target state description vector determined according to the first target state description information of the j-th first training generated image, is the first target scene description vector of the j-th first training generated image output by the content category discrimination model, is the transposed vector of, is the target scene description vector determined according to the first target scene description information of the j-th first training generated image, and min is the function of taking the minimum value.

[0058] According to an embodiment of the present invention, still in the above-mentioned adversarial training manner, in formula (3), if the input image is the first sample image, then in the training of the content category discrimination model, make maximize, where, is the similarity between the first target type description vector and the target type description vector of the content annotation information (error-free annotation information) of the first sample image, is the similarity between the first target state description vector and the target state description vector of the content annotation information of the first sample image, is the similarity between the first target scene description vector and the target scene description vector of the content annotation information of the first sample image. Maximizing the minimum value of the three similarities means maximizing all three similarities, that is, approaching 1, so that the content category discrimination model can accurately identify the type, state (such as actions, postures, expressions, etc.) and scene (such as the environment, location, etc.) of the target object in the first sample image. If the input image is the first training generated image, then in the training of the content category discrimination model, make minimize, that is, as close to 0 as possible, so that maximize, so that when the content category discrimination model identifies the content in the generated image, it determines that it is a generated image and does not contain any real content, thereby improving the accuracy of the content category discrimination model, where, is the similarity between the first target type description vector and the target type description vector of the first target type description information, is the similarity between the first target state description vector and the target state description vector of the first target state description information, is the similarity between the first target scene description vector and the target scene description vector of the first target scene description information. Therefore, even if the authenticity discrimination model determines that the probability of the first training generated image being a real image is not 0, the content category discrimination model can still determine that the above three similarities are close to 0, that is, it is determined that the first training generated image does not contain real content.

[0059] According to an embodiment of the present invention, in formula (3), if the input image is the first training generated image, then in the training of the intelligent image generation model, make maximized, that is, as close to 1 as possible, so that is minimized, so as to improve the values of the three similarities under the influence of the probability that the first training generated image is a real image, make the content of the first training generated image more in line with expectations, improve the performance of the intelligent image generation model, and at the same time improve the authenticity of the images generated by the intelligent image generation model.

[0060] In this way, the first content loss function can be determined through adversarial training, so that while improving the accuracy of the content generated by the intelligent image generation model, the ability of the content category discrimination model to judge whether the image content is correct can be improved, so as to balance the performance of the two models, further improve the authenticity of the first training generated image, and the accuracy of the content included, so that the content in the images generated by the intelligent image generation model is more in line with expectations.

[0061] According to an embodiment of the present invention, obtaining the combined loss function according to the first authenticity loss function, the first industry loss function, and the first content loss function includes: obtaining the combined loss function according to formula (4) ,

[0062] (4)

[0063] wherein, represents adjusting the parameters of the intelligent image generation model in the direction of minimizing , represents adjusting the parameters of the discrimination model group in the direction of maximizing , is the first authenticity loss function, is the first industry loss function, is the first content loss function.

[0064] According to an embodiment of the present invention, in formula (4), the above three loss functions can be summed up, and the trained intelligent image generation model is trained using the summed loss function. When training the intelligent image generation model, it is trained in a way that minimizes the summed loss function. When training the discriminant model group, it is trained in a way that maximizes the summed loss function, so as to balance the performance of the intelligent image generation model and multiple models in the discriminant model group, thereby improving the performance of the intelligent image generation model while enhancing the discriminant ability of the discriminant model group.

[0065] According to an embodiment of the present invention, in step S104, the intelligent image generation model and the discriminant model group can be trained multiple times in the above manner. After the performance of the intelligent image generation model and multiple models in the discriminant model group reaches balance, the training can be completed to obtain the trained intelligent image generation model.

[0066] According to an embodiment of the present invention, in step S105, the obtained intelligent image generation model that can be applied to multiple industry types can be used to train a lightweight dedicated image generation model. The number of parameters of the dedicated image generation model is less than that of the intelligent image generation model, and it is dedicated to the image generation work of a specific industry. Compared with the intelligent image generation model, its complexity is lower, and the purpose of rapid construction and training can be achieved, so as to quickly obtain a lightweight model applicable to a specific industry. The second training generated image can be obtained according to the second industry category information, the second demand description information, and the trained intelligent image generation model. For example, the description feature vector can be obtained based on the second industry category information and the second demand description information and input into the trained intelligent image generation model to obtain the second training generated image. The second training generated image can be considered as an image that conforms to the industry characteristics of the second industry category information and contains content consistent with the description of the second demand description information, and can be used as a reference for training the dedicated image generation model.

[0067] According to an embodiment of the present invention, in step S106, since the dedicated image generation model is dedicated to the industry corresponding to the second industry category information, the second industry category information does not need to be input, and only the second demand description information is input, so as to train the ability of the dedicated image generation model to output an image that conforms to the industry characteristics and can accurately express the content contained in the second demand description information. The dedicated image generation model can generate the third training generated image.

[0068] According to an embodiment of the present invention, in step S107, taking the second training generated image as a reference, the error of the third training generated image can be determined and reduced, so that the image generated by the dedicated image generation model is as consistent as possible with the image generated by the intelligent image generation model when generating images of this industry.

[0069] According to an embodiment of the present invention, the dedicated loss function of the dedicated image generation model is obtained based on the third training-generated image and the second training-generated image, including: obtaining the dedicated loss function of the dedicated image generation model according to formula (5). ,

[0070] (5)

[0071] wherein, is the pixel value of the pixel point with coordinates in the second training-generated image, is the pixel value of the pixel point with coordinates in the third training-generated image, L is the length of the second training-generated image, and W is the width of the second training-generated image.

[0072] According to an embodiment of the present invention, in formula (5), the error of the pixel values of the third training-generated image and the second training-generated image at each pixel point can be directly determined, and the sum of the errors is used as the loss function. During the training process, the loss function is minimized, so as to improve the consistency between the image generated by the dedicated image generation model and the image generated by the intelligent image generation model in this industry. When the dedicated image generation model generates images in this industry, it can also achieve the same level of authenticity, content accuracy, and rendering effect as the image generated by the intelligent image generation model.

[0073] In this way, the dedicated loss function can be determined based on the error of the pixel values of the third training-generated image and the second training-generated image at each pixel point, so as to improve the consistency between the image generated by the dedicated image generation model and the image generated by the intelligent image generation model in this industry. When the dedicated image generation model generates images in its corresponding industry, it can achieve a similar effect to the intelligent image generation model, and is more lightweight, facilitating construction and training.

[0074] According to an embodiment of the present invention, in step S108, the dedicated image generation model can be trained based on the above dedicated loss function, and the trained dedicated image generation model can be obtained after multiple trainings, so that the trained dedicated image generation model can achieve a similar effect to the intelligent image generation model when generating images in the corresponding industry, and realize fast construction and training.

[0075] The method for constructing an intelligent image generation model for diversified industry applications according to an embodiment of the present invention can train the intelligent image generation model through first industry category information and first requirement description information of multiple industry categories, enabling the intelligent image generation model to generate images that meet industry applications, facilitating use in diversified industry application scenarios, and can use the intelligent image generation model to train dedicated image generation models for each industry, facilitating the construction of lightweight dedicated image generation models based on industries and fields, so as to quickly construct models adapted to the usage scenarios of various industries. The first authenticity loss function can be determined through adversarial training, so as to be able to improve the authenticity discrimination ability of the authenticity discrimination model while meeting the requirement of improving the realism of the images generated by the intelligent image generation model, thereby being able to balance the performance of the two models and improve the realism of the first training generated images. The first industry loss function can be determined through adversarial training, so as to be able to improve the ability of the industry category discrimination model to judge the industry category to which the image belongs while improving the ability of the intelligent image generation model to describe industry characteristics and the accuracy of the rendering effect, thereby being able to balance the performance of the two models and further improve the authenticity of the first training generated images and the accuracy of the rendering effect. The first content loss function can be determined through adversarial training, so as to be able to improve the ability of the content category discrimination model to judge whether the content of the image is correct while improving the accuracy of the content generated by the intelligent image generation model, thereby being able to balance the performance of the two models and further improve the authenticity of the first training generated images and the accuracy of the content contained therein, making the content in the images generated by the intelligent image generation model more in line with expectations. The dedicated loss function can be determined based on the error of the pixel values of the third training generated image and the second training generated image at each pixel point, thereby improving the consistency between the images generated by the dedicated image generation model and the images of the industry generated by the intelligent image generation model, enabling the dedicated image generation model to achieve an effect similar to that of the intelligent image generation model when generating images of its corresponding industry, and being able to be more lightweight, facilitating construction and training.

[0076] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and described in the embodiments. Without departing from the above principles, the embodiments of the present invention can have any deformation or modification.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing an intelligent image generation model for diversified industry applications, characterized in that: include: Obtaining a first description feature vector according to the first industry category information and the first requirement description information, wherein the first industry category information is used to describe the industry category corresponding to the generated image, and the first requirement description information is used to describe the content contained in the generated image; Processing the first description feature vector according to the intelligent image generation model to obtain a first training generated image; Obtaining a combined loss function of the intelligent image generation model and the discriminant model group according to the first training generated image, the first sample image set, and the discriminant model group, wherein the first sample image set includes first sample images of multiple industry categories; The intelligent image generation model is trained by using the combined loss function to obtain a trained intelligent image generation model; Obtaining a second training generated image according to the second industry category information and the second demand description information, and the trained intelligent image generation model; Input the second requirement description information into a dedicated image generation model corresponding to the second industry category information to obtain a third training generated image; Obtaining a dedicated loss function of a dedicated image generation model according to the third training generated image and the second training generated image; The dedicated image generation model is trained according to the dedicated loss function to obtain a trained dedicated image generation model.

2. The method for constructing an intelligent image generation model for diversified industry applications according to claim 1, characterized in that: The discrimination model group includes an authenticity discrimination model, an industry category discrimination model and a content category discrimination model, wherein the authenticity discrimination model is used to judge the authenticity of an input image, the industry category discrimination model is used to judge the industry category corresponding to the input image, and the content category discrimination model is used to judge the category of content contained in the input image.

3. The method for constructing an intelligent image generation model for diversified industry applications according to claim 2, characterized in that: According to the first training generated image, the first sample image set, and the discriminant model group, a combined loss function of the intelligent image generation model and the discriminant model group is obtained, including: Inputting the first training generated image or the first sample image into an authenticity discrimination model to obtain a first authenticity discrimination result; Determining a first authenticity loss function of an authenticity discrimination model according to the first authenticity discrimination result; Inputting the first training generated image or the first sample image into an industry category discrimination model to obtain a first industry category description vector; Obtaining a first industry loss function of an industry category discrimination model according to the first industry category description vector, the first industry category information, and the industry category of the first sample image; Inputting the first training generated image or the first sample image into a content category discrimination model to obtain a first content description vector; Obtaining a first content loss function of a content category discrimination model according to the first content description vector, the first requirement description information, and content annotation information of the first sample image; The combined loss function is obtained according to the first authenticity loss function, the first industry loss function and the first content loss function.

4. The method for constructing an intelligent image generation model for diversified industry applications according to claim 3 is characterized in that: Determining a first authenticity loss function of an authenticity discrimination model according to the first authenticity discrimination result includes: According to the formula Determine the first authenticity loss function L of the authenticity discrimination model 1,A , where SP i is the i-th first sample image, D(SP i ) is the probability that the first sample image of the ith image output by the authenticity discrimination model is a real image, TP j Generate the image for the jth first training, D(TP j ) is the probability that the jth first training generated image output by the authenticity discrimination model is a real image, n1 is the number of first sample images in a training batch, and m1 is the number of first training generated images in a training batch.

5. The method for constructing an intelligent image generation model for diversified industry applications according to claim 4, characterized in that: According to the first industry category description vector, the first industry category information and the industry category of the first sample image, a first industry loss function of the industry category discrimination model is obtained, including: Obtaining rendering parameter vectors of a plurality of first sample images of each industry category in the first sample image set, wherein each component of the rendering parameter vector is a rendering parameter, and the rendering parameters include a light and shadow contrast parameter, a saturation parameter, a brightness parameter, and an area ratio between a target area and a background area; Determine an average rendering parameter vector for each industry category based on the rendering parameter vectors of the plurality of first sample images; Obtaining a training rendering parameter vector of a first training generated image or a first sample image input into an industry category discrimination model; A first industry loss function of an industry category discrimination model is determined according to the training rendering parameter vector, the average rendering parameter vector, the first industry category description vector and the first industry category information.

6. The method for constructing an intelligent image generation model for diversified industry applications according to claim 5, characterized in that: Determining a first industry loss function of an industry category discrimination model according to the training rendering parameter vector, the average rendering parameter vector, the first industry category description vector, and the first industry category information includes: According to the formula Determine the first industry loss function L of the industry category discrimination model 1,I , where L D (SP i ) is the first industry category description vector of the i-th first sample image output by the industry category discrimination model, [L D (SP i )] T For L D (SP i ), L AN (SP i ) is the category description vector determined according to the industry category of the i-th first sample image, L D (TP j ) is the first industry category description vector of the jth first training generated image output by the industry category discrimination model, [L D (TP j )] T For L D (TP j ), L IN (TP j ) is the category description vector determined according to the first industry category information of the jth first training generated image, L TR (TP j ) is the training rendering parameter vector of the jth first training generated image, [L TR (TP j )] T For L TR (TP j ), L R (TP j ) is the average rendering parameter vector of the first industry category for the jth first training generated image.

7. The method for constructing an intelligent image generation model for diversified industry applications according to claim 4, characterized in that: The first content description vector includes a first target type description vector, a first target state description vector and a first target scene description vector, and the first requirement description information includes first target type description information, first target state description information and first target scene description information; Obtaining a first content loss function of a content category discrimination model according to the first content description vector, the first requirement description information, and the content annotation information of the first sample image includes: According to the formula Get the first content loss function L of the content category discrimination model 1,C , where L C (SP i ) is the first target type description vector of the i-th first sample image output by the content category discrimination model, [L C (SP i )] T For L C (SP i ), L ANC (SP i ) is the target type description vector determined according to the content annotation information of the i-th first sample image, L S (SP i ) is the first target state description vector of the i-th first sample image output by the content category discrimination model, [L S (SP i )] T For L S (SP i ), L ANS (SP i ) is the target state description vector determined according to the content annotation information of the i-th first sample image, L SC (SP i ) is the first target scene description vector of the i-th first sample image output by the content category discrimination model, [L SC (SP i )] T For L SC (SP i ), L ANSC (SP i ) is the target scene description vector determined according to the content annotation information of the i-th first sample image, L C (TP j ) is the first target type description vector of the jth first training generated image output by the content category discrimination model, [L C (TP j )] T For L C (TP j ), L NC (TP j ) is the target type description vector determined according to the first target type description information of the jth first training generated image, L S (TP j ) is the first target state description vector of the jth first training generated image output by the content category discrimination model, [L S (TP j )] T For L S (TP j ), L NS (TP j ) is the target state description vector determined according to the first target state description information of the jth first training generated image, L SC (TP j ) is the first target scene description vector of the jth first training generated image output by the content category discrimination model, [L SC (TP j )] T For L SC (TP j ), L NSC (TP j ) is the target scene description vector determined according to the first target scene description information of the j-th first training generated image, and min is the minimum value function.

8. The method for constructing an intelligent image generation model for diversified industry applications according to claim 3, characterized in that: Obtaining the combined loss function according to the first authenticity loss function, the first industry loss function, and the first content loss function includes: According to the formula LOSS CO =who G max D (L 1,A +L 1,I +L 1,c ) Get the combined loss function LOSS CO , where min G Indicates that according to (L 1,A +L 1,I +L 1,C ) minimizes the direction of adjusting the parameters of the intelligent image generation model, max D Indicates that according to (L 1,A +L 1,I +L 1,C ) is maximized to adjust the parameters of the discriminant model group, L 1,A is the first authenticity loss function, L 1,I is the first industry loss function, L 1,C is the first content loss function.

9. The method for constructing an intelligent image generation model for diversified industry applications according to claim 1, characterized in that: Obtaining a dedicated loss function of a dedicated image generation model according to the third training generated image and the second training generated image, including: According to the formula Get the dedicated loss function LOSS of the dedicated image generation model SP , where p2(x, y) is the pixel value of the pixel with coordinates (x, y) in the second training generated image, p3(x, y) is the pixel value of the pixel with coordinates (x, y) in the third training generated image, L is the length of the second training generated image, and W is the width of the second training generated image.

Citation Information

Patent Citations

  • Training method of image generation model, image generation method and related equipment

    CN118114788A

  • Zero-sample image recognition method and system based on generative adversarial network

    CN111476294A

  • Image segmentation model training method and device, image segmentation method and device and electronic equipment

    CN112330685A