Material image generation method and device, electronic equipment and storage medium
By acquiring the main image of the material and using the literary and artistic image model to generate different types of material images, the problem of time-consuming and costly acquisition of advertising and marketing materials in the prior art is solved, and rapid and economical material generation is achieved.
Patent Information
- Application Number
- CN202311696932.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
The method of obtaining advertising and marketing materials in the prior art takes time and is costly, making it difficult to quickly obtain effective and high-quality materials.
By obtaining the main image of the material, determining its target color matching strategy and material template, and using the literary picture model to generate multiple different styles of material images based on the target template to meet user needs.
This reduces the large amount of labor and time costs required to acquire material images, and achieves rapid and batch generation of material images that meet user needs.
Smart Images

Figure CN120147475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a material image generation method, device, electronic equipment and storage medium. Background Art
[0002] Advertising and marketing are crucial modules for business operations. In the marketing plans of most companies, advertising always occupies a place, mainly due to its advantage of achieving precision marketing. Advertising and marketing materials are an indispensable part of it, but the construction of advertising and marketing materials requires the construction of differentiated advertising and marketing materials based on the needs of different customer groups in order to achieve the purpose of precision marketing for thousands of people. This method consumes a lot of manpower and time costs. At the same time, with the continuous soaring of labor and time costs, it is becoming increasingly difficult to obtain effective and high-quality advertising and marketing materials.
[0003] In related technologies, picture-type creative materials are usually used as advertising and marketing materials. However, picture-type creative materials are mostly expressed in real time and linkage, so they need to be obtained by artificial design: personalized design is carried out according to the marketing implementation to achieve consistency between the materials and the implementation scenarios.
[0004] The way to obtain advertising and marketing materials in related technologies basically relies on the creation of designers or original artists. Not only is the work process cumbersome and time-consuming, but it also requires a lot of manpower and time costs. Therefore, how to quickly obtain advertising and marketing materials is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The purpose of the embodiments of the present invention is to provide a material image generation method, device, electronic device and storage medium to solve the technical problems of long time consumption and high cost in the related art in obtaining advertising and marketing materials.
[0006] In a first aspect, an embodiment of the present invention provides a material image generation method, comprising:
[0007] Get the main image of the material;
[0008] Determining a target color matching strategy for the main material image according to the main material image;
[0009] Determining a target material template for the main material image according to the main material image, the target color matching strategy, and target scene information, wherein the target scene information is scene information of an application scene corresponding to the main material image;
[0010] Use the first requirement information that the material image needs to meet and the target scene information as prompt words, and input them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the material main image.
[0011] In a second aspect, an embodiment of the present invention provides a material image generation device, including:
[0012] An acquisition module, configured to acquire a material main image;
[0013] A first determination module, configured to determine a target color matching strategy for the material main image according to the material main image;
[0014] A second determination module, configured to determine a target material template for the material main image according to the material main image, the target color matching strategy, and target scene information, where the target scene information is the scene information of the application scene corresponding to the material main image;
[0015] A first generation module, configured to use the first requirement information that the material image needs to meet and the target scene information as prompt words, and input them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the material main image.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the material image generation method described in any one of the above.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the material image generation method described in any one of the above.
[0018] An embodiment of the present invention provides a material image generation method, device, electronic device, and storage medium. By acquiring the material main image, this method can determine the target color matching strategy that matches the material main image and combine the scene information corresponding to the material main image, and can generate a corresponding material template for the material main image, so as to facilitate the use of a text-to-image large model to batch generate multiple different styles of target material images that meet user requirements on the basis of the material template, effectively reducing the large amount of labor cost and time cost required to obtain material images. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flowchart of a material image generation method provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic flowchart of a method for obtaining a material template provided by an embodiment of the present invention;
[0021] Figure 3 It is another schematic flowchart of a method for generating a material image provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic structural diagram of a device for generating a material image provided by an embodiment of the present invention;
[0023] Figure 5 It is another schematic structural diagram of a device for generating a material image provided by an embodiment of the present invention;
[0024] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention;
[0025] Figure 7 It is another schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0028] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0029] In the related art, picture-based creative materials are usually used as advertising and marketing materials, and picture-based creative materials are mostly expressed in real-time and linkage, so they need to be obtained by manual design: personalized design is carried out according to the marketing landing to achieve the consistency between the materials and the landing scenario.
[0030] The way to obtain advertising and marketing materials in related technologies basically relies on the creation of designers or original artists. Not only is the work process cumbersome and time-consuming, but it also requires a lot of manpower and time costs. Therefore, how to quickly obtain advertising and marketing materials is a technical problem that needs to be solved urgently.
[0031] In order to solve the technical problems existing in the related art, the embodiment of the present invention provides a material image generation method, see Figure 1 , Figure 1 is a flow chart of a material image generation method provided by an embodiment of the present invention, the method comprising steps 101 to 104;
[0032] Step 101, obtaining the main image of the material.
[0033] In this embodiment, the material main image provided in this embodiment is an image containing only the main object part and not the background part, that is, the image input by the user itself is a processed image. Therefore, by using an image containing only the main object part, a material image for the material main image can be generated when a material image is subsequently generated, effectively improving the accuracy of material image generation.
[0034] In some embodiments, the image input by the user may be an image containing a background portion, that is, the image input by the user is an unprocessed image. Therefore, when the present embodiment detects that the initial image input by the user is an image containing a background portion, that is, unprocessed, it is necessary to extract the main object in the initial image to obtain the main material image. Specifically, the steps of obtaining the main material image provided by the present embodiment may be: obtaining the initial image to be processed; identifying and segmenting the main object in the initial image to obtain the main material image in the initial image.
[0035] In this embodiment, the recognition and segmentation processing adopted in this embodiment includes recognition processing and segmentation processing, wherein the recognition processing provided in this embodiment may be recognition processing performed by using a recognition algorithm such as the YOLO-V5 algorithm, and the segmentation processing provided in this embodiment may be a segmentation algorithm performed by using a segmentation algorithm such as the segment anything algorithm. Specifically, the recognition algorithm provided in this embodiment mainly recognizes the initial image to output the position of the main object in the initial image; the segmentation algorithm provided in this embodiment mainly segments the main object in the initial image from the initial image to obtain a main image of the material without the background part.
[0036] It should be noted that the recognition algorithm and segmentation algorithm provided in this embodiment are not limited to the algorithms mentioned above, and other algorithms that can also perform recognition and segmentation processing are also available, which are not listed one by one here.
[0037] Step 102: Determine the target color matching strategy for the main image of the material according to the main image of the material.
[0038] In this embodiment, since the main image of the material provided in this embodiment is an image that only contains the main object part, that is, the main image of the material does not contain any background part, therefore, this embodiment needs to select a matching target color matching strategy for the main image of the material to determine the color of the background part that matches the main image of the material, so as to achieve the purpose of improving the performance effect of the subsequent generated material image.
[0039] Specifically, in order to effectively improve the performance effect of the subsequent generated material image, this embodiment can use a pre-trained target classification model to classify the main image of the material to determine the target color matching strategy that matches the main image of the material. In order to obtain the target classification model, the method for generating a material image provided in this embodiment may further include, before the step of determining the target color matching strategy for the main image of the material according to the main image of the material: using a large number of training pair data to perform classification training on the classification model to be trained until convergence to obtain the target classification model.
[0040] Among them, the training pair data provided in this embodiment includes the training main image of the material and the color matching strategy corresponding to the training main image of the material. Specifically, the classification of the color matching strategy provided in this embodiment, that is, the correspondence between the training main image of the material and the color matching strategy in the training pair data, is determined based on an expert knowledge base, and this expert knowledge base contains all knowledge in the field of advertising and marketing, so as to provide high-quality color matching strategies for different training main images of the material.
[0041] In this embodiment, the expert knowledge base provided in this embodiment can perform color matching on different categories of training main images of the material from the industry perspective, performance characteristics perspective, and other professional perspectives to obtain matching color matching strategies:
[0042] From the industry perspective, the training main image of the food category can use light yellow and pink to give people a feeling of warmth and closeness; the training main image of the beverage category can use green and blue; the training main image of the wine and pastry categories can use bright red; the training main image of the daily cosmetics category can use rose color, pinkish white, light green, light blue, and dark coffee color to highlight the warm and elegant sentiment; the training main image of the clothing and footwear category can use dark green, dark blue, coffee color or gray to highlight the beauty of solemnity and elegance.
[0043] In terms of performance characteristics, just looking at the main images of food training materials, cakes and desserts mostly use gold, yellow, and light yellow to give people a fragrant impression; beverages such as tea and beer mostly use red or green, symbolizing the richness and aroma of tea; tomato juice and apple juice mostly use red, which concentrates on indicating the natural attributes of the item.
[0044] From other professional perspectives, for example, when the main image of the training material is a light blue mobile phone, its color matching strategy can be closer to a blue-green background. A simple background can create a gentle, elegant, and high-end feeling. At the same time, the color matching strategy provided by this embodiment can also include light and dark contrast; cold and warm contrast, such as the contrast between red and blue; dynamic and static contrast, such as the contrast between a light and quiet background and lively and chaotic patterns and text; light and heavy contrast, such as the contrast between deep pigments and light pigments, etc.
[0045] Thus, the color matching strategy provided by the embodiment of the present invention can effectively improve the performance of the material image, thereby improving the marketing effect of using the material image as a marketing advertisement. Therefore, the color matching strategy provided by this embodiment can obtain high-quality training data.
[0046] In this embodiment, by adopting the high-quality training provided by this embodiment to classify the data for the classification model to be trained, the classification model can learn the knowledge in the expert knowledge base, so that the target classification model after the training has the ability to classify color matching strategies for different material subject images, reducing the need for professional designers to perform long-term color matching strategy configuration processes, saving a lot of manpower costs and time costs.
[0047] After obtaining the target classification model, the step of determining the target color matching strategy of the main material image according to the main material image provided in this embodiment can be: inputting the main material image into the target classification model to obtain the target color matching strategy corresponding to the main material image. In this way, using the target classification model provided in this embodiment, different main material images can be classified, thereby achieving the purpose of providing high-quality and matching color matching strategies for different main material images.
[0048] Step 103: determining a target material template of the main material image according to the main material image, the target color matching strategy and the target scene information.
[0049] The target scene information provided in this embodiment is the scene information of the application scene corresponding to the main image of the material. Specifically, the scene information can be the category information of the category to which the main image of the material belongs, for example, the application scene of the main image of the food material is the food scene information, and the application scene of the main image of the clothing material is the clothing scene information.
[0050] In some embodiments, in order to improve the quality of the target material template, this embodiment can also adopt the expert knowledge base provided in the above embodiment to configure high-quality material templates for different material main images, so as to improve the quality of the subsequent generated material images. Specifically, please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for obtaining a material template provided by an embodiment of the present invention. As Figure 2 shown, the method for obtaining a material template provided in this embodiment includes steps 201 to 203;
[0051] Step 201, determine the target configuration requirement information of the material main image according to the material main image, the target color matching strategy, and the target scene information.
[0052] In this embodiment, this embodiment can not only be based on the material theme image, the target color matching strategy, and the target scene information, but also be based on the user group information input by the user and the user's requirement information, and use this as the target configuration requirement information of the material main image.
[0053] Among them, the more target configuration requirement information provided for the material main image, that is, the more detailed the target configuration requirement information, the more it can improve the quality of the subsequent obtained target material template. Therefore, the target configuration requirement information provided in this embodiment is not limited to the data mentioned in the above embodiment, and can also be other data that can improve the quality of the subsequent obtained target material template, which is not specifically limited here.
[0054] Step 202, based on the target configuration requirement information, perform a matching process in the target mapping relationship table to obtain a matching result.
[0055] Among them, the target mapping relationship table provided in this embodiment includes multiple configuration requirement information and the corresponding material templates. Specifically, the target mapping relationship table provided in this embodiment is constructed based on the expert knowledge base provided in the above embodiment. In this expert knowledge base, high-quality material templates are configured for different configuration requirement information. Using this high-quality material template can effectively prompt the subsequent text-to-image model to generate high-quality material images.
[0056] Optionally, the matching process provided in this embodiment can be a similarity calculation process, that is, calculate the similarity between the target configuration requirement information and each configuration requirement information in the target mapping relationship table, so as to obtain the matching result, that is, the similarity value.
[0057] Step 203, use the material template corresponding to the configuration requirement information whose matching result meets the preset conditions as the target material template of the material main image.
[0058] In this embodiment, the preset condition provided in this embodiment may be that the value of the matching result is greater than a preset value, such as 80%. For example, in the target mapping relationship table, the high-quality material template corresponding to the configuration requirement information with a matching result (similarity value) greater than 80% of the target configuration requirement information is determined as the target material template.
[0059] In some embodiments, when there are multiple matching results that meet the preset conditions, the material template corresponding to the configuration requirement information with the optimal matching result can be selected as the target material template. For example, the material template corresponding to the configuration requirement information with a similarity value greater than the preset condition of 80% and the largest is used as the target material template. It should be noted that the number of target material templates provided in this embodiment is not limited to one, and can also be multiple. The specific value can be set according to actual application requirements, so the number of target material templates is not specifically limited here.
[0060] Step 104: Use the first requirement information that the material image needs to meet and the target scene information as prompt words, and input them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the material main image.
[0061] Among them, the first requirement information provided in this embodiment is the requirement information input by the user, which the target text-to-image model is expected to generate the target material image to meet. The target scene information can be input by the user, or the pre-trained recognition model can be used to classify and recognize the material main image to determine the category information of the category to which the material main image belongs. By adding scene information to the prompt words, the target text-to-image model can better understand the current generation task, effectively improving the accuracy of the target text-to-image model in generating images.
[0062] Specifically, the first requirement information may include information such as user-defined requirement information and user group information, as long as it is prompt information used to prompt the target text-to-image model to generate a material image, and no specific limitation is made here.
[0063] As an optional embodiment, in order to enable the target text-to-image model to generate a target material image with better effects and higher quality, this embodiment can also perform instruction fine-tuning on a large language model, such as the GPT large language model, so that the large language model generates natural language corresponding to the prompt words that can enable the target text-to-image model to better understand the problem and intention based on the first requirement information and the target scene information. Then, the natural language output by the large language model is used as a prompt word and input into the target text-to-image model together with the target material template, thereby helping the target text-to-image model better understand the problem and intention, and enabling the target text-to-image model to generate a target material image with better effects and higher quality.
[0064] Specifically, the method for fine-tuning the large language model provided in this embodiment can be as follows:
[0065] The first demand information and the target scenario information + "Please help me edit a prompt for input into the text-to-image model based on the first demand information and the target scenario information to help the target text-to-image model better understand the problem and intention."
[0066] After fine-tuning the large language model with the above instructions, the natural language corresponding to the prompt that can enable the target text-to-image model to better understand the problem and intention can be obtained from the feedback of the large language model.
[0067] At the same time, the user can also fine-tune the large language model so that the large language model outputs a prompt that can enable the target text-to-image model to generate more personalized target material images, thereby realizing the customization of the content of the material images, meeting the requirements for generating specific content, and effectively improving the user experience.
[0068] In this embodiment, the target text-to-image model provided in this embodiment can be models such as Imagen, DALL-E series, Stable Diffusion, VQGAN, etc., or other models with text-to-image functions, which are not specifically limited here. Optionally, this embodiment mainly uses the Stable Diffusion model as the basic model, and improves and trains it to obtain the target text-to-image model provided in this embodiment. Specifically, this embodiment can use the Lora method to fine-tune the Stable Diffusion model to obtain a more targeted target text-to-image model that can more effectively improve the image generation effect in the marketing scenario.
[0069] The target material template provided in this embodiment can be used as an image prompt, so as to combine with the text prompt of the first demand information and the target scenario information to prompt the target text-to-image model to generate a material image that meets the first demand information and the target scenario information on the basis of the target material template. In this way, the target material image corresponding to the material main image input by the user can be quickly obtained, effectively solving the technical problem of high time cost and labor cost for obtaining material images in the related art.
[0070] Optionally, since the target text-to-image model provided in this embodiment will generate multiple target material images with different styles when generating the target material image, the user can select the target material image they need from the multiple target material images. In this way, by generating a batch of target material images for the user to select, the user experience can be effectively improved.
[0071] So far, the introduction and description of the material image generation method provided in this embodiment are completed.
[0072] However, as an alternative embodiment, in order to further improve the marketing effect of the target material image for advertising and marketing, the method for generating a material image provided in this embodiment may further include: inputting the second requirement information that the material copy needs to meet and the target material image into a large language model to obtain the target material copy returned by the large language model.
[0073] After obtaining the final target material image, this embodiment can also access a large language model, such as the GPT large language model, and perform instruction fine-tuning on the large language model so that the large language model provides creative and highly immersive material copy for advertising and marketing. By combining the material image with the material copy, it is possible to successfully evoke an emotional resonance while also helping to strengthen the image of the marketed product.
[0074] It should be noted that the large language model provided in this embodiment is not limited to the GPT large language model mentioned in the above embodiment, and GPT4 can also be selected as the large language model, as long as it can generate the material copy corresponding to the target material image, and no specific limitation is made here.
[0075] In this embodiment, the method for performing instruction fine-tuning on the large language model provided in this embodiment can be as follows:
[0076] Target material picture + "If you are a marketing material copy advertisement designer for clothing products, please help me write a professional marketing material copy for the above marketing material."
[0077] After fine-tuning the large language model through the above instructions, the material copy corresponding to the target material image feedback by the large language model can be obtained.
[0078] In some embodiments, in order to further improve the quality of the generated material image to improve the final marketing effect for advertising and marketing, please refer to Figure 3 , Figure 3 is another flow schematic diagram of the method for generating a material image provided by an embodiment of the present invention. As shown in Figure 3 , the method for generating a material image provided in this embodiment includes steps 301 to 308;
[0079] Step 301, obtain a material main image.
[0080] In this embodiment, the material main image provided in this embodiment is an image that only contains the main object part and does not contain the background part. That is, the image input by the user is an image that has been processed. Therefore, by using an image that only contains the main object part, it is possible to generate a material image for the material main image during subsequent generation of the material image, effectively improving the accuracy of material image generation.
[0081] When the input image of the user is unprocessed, that is, the content contains the background part, it can be recognized and segmented to obtain a material main image that only contains the main object part. For the specific recognition and segmentation processing, please refer to the method provided in the above embodiments, which will not be elaborated here.
[0082] Step 302: Determine the target color matching strategy of the material main image according to the material main image.
[0083] In this embodiment, this embodiment can use a pre-trained target classification model to classify the material main image, and then select a matching target color matching strategy according to the finally classified category. For the specific classification processing and the method of selecting a matching target color matching strategy, please refer to the method provided in the above embodiments, which will not be elaborated here.
[0084] Step 303: Determine the target material template of the material main image according to the material main image, the target color matching strategy and the target scene information.
[0085] Among them, the target scene information is the scene information of the application scene corresponding to the material main image.
[0086] In this embodiment, this embodiment can use a preset target mapping relationship table to determine the target material template matched by the material main image. For the specific determination method, please refer to the method provided in the above embodiments, which will not be elaborated here.
[0087] Step 304: Determine the target application scene where the material main image is located according to the target scene information.
[0088] In this embodiment, since the target scene information provided in this embodiment is the scene information of the application scene corresponding to the material main image, therefore, the target application scene where the material main image is located can be directly determined according to the target scene information. For example, when the material main image is clothing, the target application scene where it is located is the clothing application scene.
[0089] Specifically, the target application scene provided in this embodiment is not limited to the clothing application scene, and can also be other application scenes such as food and daily necessities, which will not be specifically limited here.
[0090] Step 305: Select a text-to-image generation model that matches the target application scene from a preset multiple text-to-image generation models as the target text-to-image generation model.
[0091] In this embodiment, multiple text-to-image models can be preset, such as text-to-image model A, text-to-image model B, and text-to-image model C. Among them, different text-to-image models are trained with training data under different application scenarios, so that different text-to-image models can generate more targeted material images under different application scenarios.
[0092] For example, text-to-image model A is mainly for the clothing application scenario, text-to-image model B is mainly for the food application scenario, and text-to-image model C is mainly for the daily necessities application scenario. When the target application scenario where the main material image input by the user is located is the clothing application scenario, then text-to-image model A is used as the target text-to-image model matching the target application scenario; when the target application scenario where the main material image input by the user is located is the food application scenario, then text-to-image model B is used as the target text-to-image model matching the target application scenario; when the target application scenario where the main material image input by the user is located is the daily necessities application scenario, then text-to-image model C is used as the target text-to-image model matching the target application scenario. In this way, by adopting the embodiment of the present invention, a more targeted text-to-image model can be selected for image generation according to the target application scenario where different main material images are located, and the effect of text-to-image model image generation under different application scenarios can be improved more pertinently.
[0093] Step 306: Use the first requirement information and target scenario information that the material image needs to meet as prompt words, and input them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the main material image.
[0094] After selecting the target text-to-image model that matches the target main material, the first requirement information and target scenario information that the material image needs to meet can be used as prompt words, and input them together with the target material template into the target text-to-image model, so as to guide the target text-to-image model to generate high-quality target material images.
[0095] As an optional embodiment, the text-to-image large model adopted in this embodiment is constructed based on the Stable Diffusion model. Due to the randomness of the Stable Diffusion model itself, the generated material images also have uncertainty, which is the reason for different styles. Therefore, when a large number of material images need to be batch processed, due to the large amount of data for batch image generation, manually screening suitable material images will be time-consuming and inefficient, so it is not very suitable to manually screen suitable material images.
[0096] Therefore, in order to improve the screening efficiency of material images and provide users with target material images of better quality, this embodiment also provides steps 307 to 308;
[0097] Step 307, when the number of target material images is multiple, each target material image is input into the target evaluation model respectively to perform evaluation processing on the target material images, and the evaluation value of each material image is obtained.
[0098] In this embodiment, the target evaluation model provided in this embodiment is mainly used to imitate human aesthetics and objectively evaluate each target material image, so as to achieve the purpose of quickly screening out suitable and higher-quality material pictures.
[0099] Specifically, the target evaluation model provided in this embodiment can be modeled from the perspective of business, focusing on material images, color matching strategies, marketing scenarios, and picture aesthetics. The overall modeling idea can be: input training material images into the evaluation model to be trained (which can be a CNN model), extract relevant features of the material images through the evaluation model to be trained, and establish a machine learning / deep learning model to regress the aesthetic performance of the material. Then, use experts to annotate and manually evaluate and score the training material images, and let the evaluation model to be trained learn from the annotated training material images, so that the evaluation model to be trained learns the aesthetic judgment of the material images by experts, thereby obtaining the target evaluation model provided in this embodiment.
[0100] Step 308, the target material images with evaluation values greater than the preset threshold are used as preferred material images.
[0101] After obtaining the target evaluation model with the ability to judge the aesthetics of material images, the target evaluation model can be used to evaluate each target material image, so as to obtain the aesthetic evaluation value of each target material image. Subsequently, the target material images with evaluation values greater than the preset threshold can be used as preferred material images, so as to achieve the purpose of quickly screening out suitable and higher-quality preferred material images for users.
[0102] Among them, the preset threshold provided in this embodiment can be 80, 85 or 90, and the specific value of the preset threshold can be set according to actual needs and is not specifically limited here.
[0103] In summary, the embodiments of the present invention provide a method for generating a material image. The method includes obtaining a material main image, determining a target color matching strategy for the material main image according to the material main image, determining a target material template for the material main image according to the material main image, the target color matching strategy, and target scene information, where the target scene information is the scene information of the application scene corresponding to the material main image, using the first requirement information that the material image needs to meet and the target scene information as prompt words, and inputting them together with the target material template into a target text-to-image model to obtain a target material image corresponding to the material main image. By adopting the embodiments of the present invention, material images that can meet user requirements can be generated quickly and in large quantities, reducing the large amount of labor costs and time costs required to obtain material images.
[0104] According to the method described in the above embodiments, this embodiment will be further described from the perspective of a material image generation device. The material image generation device can be specifically implemented as an independent entity or integrated in an electronic device, such as a terminal. The terminal can include a mobile phone, a tablet computer, etc.
[0105] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a material image generation device provided by the embodiments of the present invention. As Figure 4 shown, the material image generation device 400 provided by the embodiments of the present invention includes: an acquisition module 401, a first determination module 402, a second determination module 403, and a first generation module 404;
[0106] Among them, the acquisition module 401 is used to acquire a material main image.
[0107] The first determination module 402 is used to determine a target color matching strategy for the material main image according to the material main image.
[0108] The second determination module 403 is used to determine a target material template for the material main image according to the material main image, the target color matching strategy, and target scene information.
[0109] Among them, the target scene information is the scene information of the application scene corresponding to the material main image.
[0110] The first generation module 404 is used to use the first requirement information that the material image needs to meet and the target scene information as prompt words, and input them together with the target material template into a target text-to-image model to obtain a target material image corresponding to the material main image.
[0111] In some embodiments, the acquisition module 401 provided in this embodiment is specifically used for: acquiring an initial image to be processed; identifying and segmenting the main object in the initial image to obtain the material main image in the initial image.
[0112] In some embodiments, the second determination module 403 provided in this embodiment is specifically configured to: determine the target configuration requirement information of the material main image according to the material main image, the target color matching strategy, and the target scene information; based on the target configuration requirement information, perform a matching process in the target mapping table to obtain a matching result, where the target mapping table includes multiple configuration requirement information and the material templates corresponding to the configuration requirement information; use the material template corresponding to the configuration requirement information whose matching result meets the preset conditions as the target material template of the material main image.
[0113] In some embodiments, please refer to Figure 5 , Figure 5 is another structural schematic diagram of the material image generation device provided by the embodiment of the present invention. As Figure 5 shown, the material image generation device 400 provided in this embodiment further includes: a second generation module 405, an evaluation module 406, a first training module 407, a third determination module 408, and a selection module 409;
[0114] Among them, the second generation module 405 is configured to input the second requirement information that the material copy needs to meet and the target material image into the large language model to obtain the target material copy returned by the large language model.
[0115] The evaluation module 406 is configured to, when the number of target material images is multiple, input each target material image into the target evaluation model respectively to perform an evaluation process on the target material images to obtain the evaluation value of each material image; use the target material image whose evaluation value is greater than the preset threshold as the preferred material image.
[0116] The first training module 407 is configured to perform classification training on the classification model to be trained using a large amount of training data until convergence to obtain the target classification model.
[0117] Among them, the training data includes training material main images and the color matching strategies corresponding to the training material main images.
[0118] In this embodiment, the first determination module 402 provided in this embodiment is specifically configured to: input the material main image into the target classification model to obtain the target color matching strategy corresponding to the material main image.
[0119] The third determination module 408 is configured to determine the target application scenario where the material main image is located according to the target scene information.
[0120] The selection module 409 is configured to select the text-to-image model that matches the target application scenario from a preset plurality of text-to-image models as the target text-to-image model.
[0121] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, please refer to the foregoing method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.
[0122] In addition, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device can be a mobile terminal such as a smart phone, a tablet computer, or the like. As Figure 6 shown, the electronic device 600 includes a processor 601 and a memory 602. Among them, the processor 601 is electrically connected to the memory 602.
[0123] The processor 601 is the control center of the electronic device 600, connects various parts of the entire electronic device through various interfaces and lines, and by running or loading application programs stored in the memory 602, and calling data stored in the memory 602, executes various functions of the electronic device 600 and processes data, thereby monitoring the electronic device 600 as a whole.
[0124] In this embodiment, the processor 601 in the electronic device 600 will load instructions corresponding to the processes of one or more application programs into the memory 602 according to the following steps, and the processor 601 will run the application programs stored in the memory 602, so as to implement any step in the material image generation method provided in the foregoing embodiment.
[0125] The electronic device 600 can implement the steps in any embodiment of the material image generation method provided by the embodiment of the present invention. Therefore, it can achieve the beneficial effects that any material image generation method provided by the embodiment of the present invention can achieve. For details, please refer to the foregoing embodiments, which will not be elaborated herein.
[0126] Please refer to Figure 7 , Figure 7 which is another schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, Figure 7 shows a specific structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device can be used to implement the material image generation method provided in the foregoing embodiment. The electronic device 700 can be a mobile terminal such as a smart phone or a laptop computer.
[0127] The RF circuit 710 is used to receive and transmit electromagnetic waves, enabling the mutual conversion between electromagnetic waves and electrical signals, thereby communicating with a communication network or other devices. The RF circuit 710 may include various existing circuit components for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 710 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through a wireless network. The aforementioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The aforementioned wireless networks can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not yet been developed currently.
[0128] The memory 720 can be used to store software programs and modules, such as the program instructions / modules corresponding to the material image generation method in the above embodiments. The processor 780 executes various functional applications and the generation of material images by running the software programs and modules stored in the memory 720.
[0129] The memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 720 may further include a memory remotely located with respect to the processor 780, and these remote memories may be connected to the electronic device 700 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0130] The input unit 730 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display screen or a touchpad, can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch-sensitive surface 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 731 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute the command sent by the processor 780. In addition, multiple types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface 731. In addition to the touch-sensitive surface 731, the input unit 730 may further include other input devices 732. Specifically, the other input devices 732 may include but are not limited to one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), trackballs, mice, joysticks, etc.
[0131] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device 700. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode). Further, the touch-sensitive surface 731 can cover the display panel 741. When the touch-sensitive surface 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in the figure, the touch-sensitive surface 731 and the display panel 741 are implemented as two independent components to perform input and output functions, in some embodiments, the touch-sensitive surface 731 and the display panel 741 can be integrated to implement input and output functions.
[0132] The electronic device 700 may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can generate an interruption when the flip cover is closed or opened. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer attitude calibration), vibration recognition-related functions (such as a pedometer, tapping), etc. As for other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, and an infrared sensor that the electronic device 700 can also be configured with, they will not be elaborated here.
[0133] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the electronic device 700. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output. On the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and converted into audio data. Then, after the audio data is output to the processor 780 for processing, it is sent to another terminal, for example, via the RF circuit 710, or the audio data is output to the memory 720 for further processing. The audio circuit 760 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 700.
[0134] The electronic device 700 can help users receive requests, send information, etc. through the transmission module 770 (such as a Wi-Fi module), which provides users with wireless broadband Internet access. Although the transmission module 770 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 700 and can be completely omitted within the scope of not changing the essence of the invention as needed.
[0135] The processor 780 is the control center of the electronic device 700, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by calling the data stored in the memory 720, it executes various functions of the electronic device 700 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 780 may include one or more processing cores; in some embodiments, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 780 either.
[0136] The electronic device 700 also includes a power supply 790 (such as a battery) for powering each component. In some embodiments, the power supply can be logically connected to the processor 780 through a power management system, thereby realizing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 790 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0137] Although not shown, the electronic device 700 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors to implement any step in the material image generation method provided in the above embodiment.
[0138] During specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the method embodiments described above, which will not be elaborated here.
[0139] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium in which multiple instructions are stored, and when these instructions are executed by a processor, any step in the material image generation method provided in the above embodiments can be implemented.
[0140] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0141] Since the instructions stored in the storage medium can execute the steps in any embodiment of the material image generation method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any material image generation method provided in the embodiments of the present invention can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0142] The above has introduced in detail a material image generation method, device, electronic device, and storage medium provided in the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application. Moreover, for those of ordinary skill in the technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A method for generating a material image, characterized in that, it includes: Obtain the main body image of the material; According to the main body image of the material, determine the target color matching strategy of the main body image of the material; According to the main body image of the material, the target color matching strategy and the target scene information, determine the target material template of the main body image of the material, where the target scene information is the scene information of the application scene corresponding to the main body image of the material; Use the first requirement information that the material image needs to meet and the target scene information as prompt words, and input them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the main body image of the material.
2. The method according to claim 1, characterized in that, the method further includes: Input the second requirement information that the material copywriting needs to meet and the target material image into the large language model to obtain the target material copywriting returned by the large language model.
3. The method according to claim 1, characterized in that, the method further includes: In the case where the number of the target material images is multiple, input each of the target material images into the target evaluation model respectively to perform evaluation processing on the target material images, and obtain the evaluation value of each material image; Use the target material images whose evaluation values are greater than the preset threshold as the preferred material images.
4. The method according to any one of claims 1-3, characterized in that, the step of obtaining the main body image of the material includes: Obtain the initial image to be processed; Identify and segment the main object in the initial image to obtain the main body image of the material in the initial image.
5. The method according to any one of claims 1-3, characterized in that, before the step of determining the target color matching strategy of the main body image of the material according to the main body image of the material, the method further includes: Use a large amount of training pair data to perform classification training on the classification model to be trained until convergence to obtain the target classification model, where the training pair data includes the training main body image of the material and the color matching strategy corresponding to the training main body image of the material; the step of determining the target color matching strategy of the main body image of the material according to the main body image of the material includes: Input the main body image of the material into the target classification model to obtain the target color matching strategy corresponding to the main body image of the material.
6. The method according to any one of claims 1-3, characterized in that, the step of determining the target material template of the main body image of the material according to the main body image of the material, the target color matching strategy and the target scene information includes: Determine the target configuration requirement information of the main body image of the material according to the main body image of the material, the target color matching strategy and the target scene information; Based on the target configuration requirement information, perform matching processing in the target mapping relationship table to obtain a matching result, where the target mapping relationship table includes multiple configuration requirement information and the material templates corresponding to the configuration requirement information; Use the material template corresponding to the configuration requirement information whose matching result meets the preset conditions as the target material template of the main body image of the material.
7. The method according to any one of claims 1-3, wherein, before the step of using the first requirement information and the target scene information that the material image needs to meet as prompt words and inputting them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the material main image, the method further includes: determining a target application scene where the material main image is located according to the target scene information; selecting, from a plurality of preset text-to-image models, a text-to-image model that matches the target application scene as the target text-to-image model.
8. A material image generation device, wherein, it includes: an acquisition module for acquiring a material main image; a first determination module for determining a target color matching strategy for the material main image according to the material main image; a second determination module for determining a target material template for the material main image according to the material main image, the target color matching strategy, and the target scene information, where the target scene information is the scene information of the application scene corresponding to the material main image; a first generation module for using the first requirement information and the target scene information that the material image needs to meet as prompt words and inputting them together with the target material template into the target text-to-image model to obtain the target material image corresponding to the material main image.
9. An electronic device, wherein, the electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the steps in the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in the method according to any one of claims 1 to 7.