Method for automatically generating picture, and electronic device
By generating semantic tags for the main object image using an AI large model and combining them with an element asset library, the problems of rapid response and professionalism in the production of marketing material images are solved, achieving efficient and professional image generation and adjustment, and applicable to diverse devices and scenarios.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2026-03-19
AI Technical Summary
Existing technologies for producing marketing material images are insufficient to meet the demands for rapid response and delivery. They are labor-intensive, slow to adjust, and difficult to apply in marketing scenarios due to the uncontrollable quality of images generated by ordinary AI.
By generating semantic tags for the main object image using a large AI model, and combining them with elements in the element asset library for vector matching and structured layout, highly professional marketing material images are generated.
It enables automated image generation, improves generation efficiency, ensures the professionalism and visual appeal of images, adapts to visual consistency across diverse devices and scenarios, and reduces labor costs.
Smart Images

Figure CN2025110393_19032026_PF_FP_ABST
Abstract
Description
Method and electronic device for automatically generating a picture
[0001] The present disclosure claims priority to Chinese Patent Application No. 202411287826.0, filed on September 13, 2024, with the Chinese Patent Office, entitled “Method and electronic device for automatically generating a picture”, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of picture generation, and in particular, to a method and electronic device for automatically generating a picture. BACKGROUND
[0003] In a commodity information service system, there is often a need to design marketing materials, which are usually in the form of pictures, and usually need to present commodity pictures, brand information, value point information, etc. in the pictures. For example, an operator in the system needs to put out a certain advertisement about the system outside the station, and needs to reflect the commodity information of high-quality commodities in the system in the advertisement link. At this time, it is necessary to design marketing material pictures based on the information of the commodities for off-site placement. Or, some merchants may need to promote their brands, and may also need to design corresponding marketing material pictures, which also need to include brand-related commodity pictures, value points, etc. information. Or, some designers may also need to design some “posters” and other marketing material pictures for promotion, etc.
[0004] In a conventional manner, a demand is usually proposed by an operator, a merchant, etc., and then a designer is responsible for designing a script, a background, a font, a layout, etc., and then the above marketing material picture can be produced. Alternatively, in order to improve efficiency, some templates can also be pre-configured, and the designer can produce the marketing material picture by replacing the goods, scripts, etc. in the template. However, in this way, the produced picture is relatively fixed and has poor maneuverability. In addition, since the designer still needs to manually replace the content of the goods, scripts, etc. in the template, the verification still depends on the manual operation of the designer. However, in actual application, especially in the scene of in-site and out-site placement by the operator, the placement rhythm is usually fast, and hundreds of pictures may need to be produced for placement every day, so multiple designers need to work continuously to catch up with the placement rhythm, which makes the human cost very high. In addition, during the placement process, it may also involve adjusting the placed picture, for example, after a large number of material pictures are placed on a certain day, it is found that the placement effect of the batch of material is not ideal, and a batch of goods may need to be replaced, or the value point presented in the material may not be attractive enough, and the script needs to be rewritten, etc. These adjustment work also needs to be done manually, so there is a problem of slow response speed, and each adjustment also means additional human cost.
[0005] In summary, in the prior art, the production of marketing material pictures based on pre-designed templates cannot meet the demand for rapid response delivery, and there is also a problem of high human cost. SUMMARY
[0006] The present disclosure provides a method and an electronic device for automatically generating pictures, which can improve the efficiency of picture generation and ensure that the automatically generated pictures have aesthetic value and professionalism in design practice.
[0007] The present disclosure provides the following solutions:
[0008] A method for generating a picture, comprising:
[0009] determining user generation demand information of a picture, the generation demand information at least comprising a subject object image;
[0010] generating a semantic label for the subject object image by an artificial intelligence (AI) large model, so as to obtain a description text for describing a feature of the subject object image according to the semantic label, and convert the description text into a first vector;
[0011] According to the first vector corresponding to the subject object image and the second vector corresponding to a plurality of different types of elements in the pre-established element asset library, a plurality of types of elements respectively matched with the subject object image are determined; wherein the element is an element required in the generation of a picture, and the second vector is obtained by converting the description text used to describe the characteristics of the element after generating a semantic label for the element by an AI large model to obtain the description text according to the semantic label.
[0012] According to the preset picture description protocol, the subject object image and the matched element are organized to generate a target picture.
[0013] Wherein, when generating a semantic label for the subject object image, the generated semantic label includes: the category of the subject object, the color tendency, the suitable occasion, the style and / or the crowd.
[0014] Wherein, when generating a semantic label for the element, the generated semantic label includes: the suitable scene of the element, the style, the crowd, the category of the subject object, the aspect ratio and / or the quantity.
[0015] Wherein, the elements in the element asset library include: layout scheme class elements, content class elements, style class elements and background class elements, wherein the content class elements include a plurality of subtypes, the subtypes include image class, text class or component class, the style class elements are used to describe the style configuration information of text and / or components, and the style includes font and / or color matching;
[0016] Wherein, the layout scheme class elements are used to define the canvas arrangement format skeleton structure, and the relative position of the content class elements in the canvas is defined by a structured relative positioning design method, so as to fill the matched content class elements into the corresponding positions in the canvas according to the definition of the matched target layout scheme class elements.
[0017] Wherein, the layout scheme class elements are used to divide the canvas into a plurality of information regions, and the subtypes of the content class elements required to be filled in each information region and the alignment manner information of the content class elements in the information region are defined, so as to fill the content class elements of the corresponding subtypes into the corresponding information regions according to the alignment manner.
[0018] Wherein, the alignment manner information of the content class elements in the information region includes:
[0019] The alignment manner of the content class elements in the main axis direction and / or the auxiliary axis direction of the information region, the main axis direction being the arrangement direction of the content class elements in the information region, and the auxiliary axis direction being the vertical direction of the main axis direction.
[0020] In the picture generation process according to the generation requirement information, for the content of the component class, the size of the background picture frame required to carry the text content and its font size information displayed in the component is determined first, and then the background picture frame of the component is drawn, so that the size of the background picture frame of the component can be adaptively changed with the text displayed in the component.
[0021] A plurality of breakpoints are set in advance in the size and / or aspect ratio dimension, and corresponding layout scheme class elements are configured in the element asset library for the plurality of breakpoints.
[0022] In the picture generation process according to the generation requirement information, the target size and / or aspect ratio of the picture to be generated are determined, and if the target size and / or aspect ratio does not hit the breakpoint, the layout scheme class element with the closest size and / or aspect ratio is selected, and the position of the content class element is adaptively determined according to the structured relative positioning design information defined by the layout scheme class element.
[0023] In the process of determining the matching element, after the matching background element is determined, the color matching class element is determined by programmatic generation.
[0024] In the process of determining the color matching class element by programmatic generation, it includes:
[0025] The main color of the matched background class element is determined, and a set of offsets is defined from the main color as the starting point. The generated main color system is changed in hue, S saturation and L lightness by the offsets, and the generated color version color meets the defined requirements.
[0026] The generation requirement information further includes text content, which includes text content to be added to the generated picture, and / or text information describing the required scene, style, and required target population.
[0027] The main object image includes a product image, and the generated target picture includes a marketing material picture for participating in a marketing activity.
[0028] A method for generating a picture, comprising:
[0029] Receiving user generation requirement information of a picture, the generation requirement information at least including a main object image;
[0030] The generated requirement information is submitted to a server, and the server is configured to generate semantic labels for the subject object image by using an AI large model, so as to obtain description text for describing features of the subject object image according to the semantic labels, and then convert the description text into a first vector; according to the first vector corresponding to the subject object image and a second vector corresponding to a plurality of different types of elements in a pre-established element asset library, a plurality of types of elements respectively matched with the subject object image are determined, and the subject object image and the matched elements are organized according to a preset picture description protocol to generate a target picture; wherein the element is an element required in the generation of the picture, and the second vector is obtained by converting description text for describing features of the element according to semantic labels generated for the element by using an AI large model;
[0031] The generated target picture is received and displayed.
[0032] A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the method of any of the preceding embodiments.
[0033] An electronic device, comprising:
[0034] One or more processors; and
[0035] A memory associated with the one or more processors, the memory configured to store program instructions that, when executed by the one or more processors, perform the steps of the method of any of the preceding embodiments.
[0036] A computer program product comprising computer program / computer executable instructions that, when executed by a processor in an electronic device, implement the steps of the method of any of the preceding embodiments.
[0037] According to the specific embodiments provided by the present disclosure, the present disclosure discloses the following technical effects:
[0038] According to the embodiment of the present disclosure, an element asset library can be established in advance for the user's demand for generating a picture, which can include various types of design elements. For each specific element instance, an AI large model can be used to generate a semantic label for the element to obtain a description text for describing the features of the element, and then the description text can be converted into a second vector for mathematical expression. After the user inputs a specific generation requirement, the AI large model can also be used to generate a semantic label for the subject object image specified by the user to obtain a description text for describing the features of the subject object image, and the description text can be converted into a first vector. In this way, the elements of various types that match the subject object image can be determined according to the first vector corresponding to the subject object image and the second vectors respectively corresponding to the various types of elements. Then, the subject object image and the matching elements can be organized according to a preset picture description protocol to generate a target picture. In this way, not only the automatic generation of the picture is realized, but also the semantic analysis of the ergonomics is applied to the picture intelligent layout scene. The method produces an intelligent visual matching and organizing logic, which can guarantee the aesthetic value and professionalism of the automatically generated picture in design practice, thereby helping to improve the visual appeal of the picture to the user and even improve the conversion rate and other indicators, and providing strong support for the successful implementation of marketing strategies.
[0039] In a preferred implementation, on the basis of the semantic labeling of the multi-modal language model and the matching of the ergonomics elements to realize the automatic production of design materials, the structured layout mode of flexible layout (Flex) and the responsive design component can be combined to realize the design rule set based on the summary of the design method, and the AI large model logic thinking ability is interspersed in it, which is helpful to realize the intelligent dynamic adjustment under various sizes and ensure that the user can provide consistent visual experience in diversified devices and scenes.
[0040] Of course, implementing any product of the present disclosure does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description only constitute some embodiments of the present disclosure, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0042] FIG. 1 is a schematic diagram of a system architecture provided by an embodiment of the present disclosure;
[0043] FIG. 2 is a flowchart of a method provided by an embodiment of the present disclosure;
[0044] FIG. 3 is a schematic diagram of a canvas information region division manner provided by an embodiment of the present disclosure;
[0045] FIG. 4 is a schematic diagram of a display effect of a content type element in an information region under a kind of alignment manner provided by an embodiment of the present disclosure;
[0046] FIG. 5 is a schematic diagram of a display effect of a content type element in an information region under another kind of alignment manner provided by an embodiment of the present disclosure;
[0047] FIG. 6 is a schematic diagram of a vector matching process provided by an embodiment of the present disclosure;
[0048] FIG. 7 is a schematic diagram of a plurality of picture generation results provided by an embodiment of the present disclosure;
[0049] FIG. 8 is a flowchart of a client-side method provided by an embodiment of the present disclosure;
[0050] FIG. 9 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present disclosure.
[0052] In the embodiments of the present disclosure, for the demand of generating pictures (including the demand of generating marketing material pictures in an e-commerce scenario, or also can be the demand of picture generation in other scenarios), the automatic generation of pictures can be performed through the capability of an AI large-scale parameter model (referred to as "AI large model") to improve efficiency and reduce or eliminate the dependence on designers and human costs. The so-called AI large model refers to a deep learning model containing a large number of parameters. Due to its large size, such AI large model can store and process a large amount of information, thereby achieving higher performance on various tasks, having strong natural language text understanding and logical thinking ability, and usually also having the processing ability of multi-modal information and content generation ability.
[0053] In a specific implementation, the application of AI large models for content generation is quite extensive. In the scenario of the embodiments of the present disclosure, a relatively simple application of AI large models can be to directly organize the demand information input by the user into a prompt text and input it into the AI large model to generate a picture. However, the quality of the picture generated in this way is usually uncontrollable. If it is a general leisure and entertainment scenario, it can be acceptable, but in marketing material deployment scenarios, the professionalism of the generated picture usually has very high requirements. Therefore, in the above-mentioned way, the picture generated by the AI large model is difficult to be directly used in the marketing scenario for on- and off-site deployment.
[0054] Therefore, in order to more effectively apply the generation capability of AI large models to scenarios with higher professional requirements, the embodiments of the present disclosure provide corresponding solutions. In this solution, an element asset library can be provided first. This element asset library can save a variety of different types of elements, which can be used in the picture generation process. For example, the types of elements can include background, font, component, color matching, copy, layout, etc. For each type, a plurality of different elements can be provided and saved in the element asset library. For example, it can include a plurality of background pictures, a plurality of font screenshots, a plurality of color matching schemes, a plurality of layout schemes, etc. Among them, the elements in the element asset library can be system-preset general assets that can be shared by multiple different users, or the user can also configure his own dedicated element asset library for use when generating pictures.
[0055] For each specific element in the element asset library, an AI large model can also generate a semantic label for the specific element to obtain a description text for describing the characteristics of the element. In addition, this description text can be converted into a vector for subsequent matching calculation. That is, for each new element added to the specific element asset library, an AI large model can be used to label it. The labeling process can be completed using structured text, thereby generating a structured description text for describing the characteristics of the specific element. Then, this description text can be converted into a vector for mathematical expression. The label of the element can include the scene, style, crowd, category of subject object, aspect ratio, and / or number of the element suitable for, etc.
[0056] In the case of establishing the above-mentioned element asset library, the user can initiate a specific picture generation requirement, and at least needs to specify a specific subject object image when initiating the requirement, which is the image that needs to be displayed as the subject in the generated picture. For example, in the product marketing scenario, the specific subject object image is usually a product image, which can be a product image of one or more products, etc. After receiving the user's requirement information, the subject object image can also be tagged by the AI large model, that is, semantic labels are added. The specific semantic labels can include the category of the subject object, color tendency, suitable occasion, style and / or people, etc. After completing the addition of semantic labels, the multiple semantic labels corresponding to the subject object image can be organized together to obtain a description text for describing the characteristics of the elements, and then the description text can also be converted into a mathematically expressed vector.
[0057] It should be noted that whether the elements are tagged or the user input subject object image is tagged, the specific tagging task can be completed by using structured text, that is, each semantic label can correspond to a structured text, for example, it can be a key-value pair structure, such as (suitable style: warm), etc. In this way, for the same element or subject object image, the structured description text can be composed according to the structured expression of the semantic label, for example, a certain subject object image has the following labels (suitable style: warm), (product color tendency: orange), (suitable background type: hand-drawn flowers), etc. The description text of the subject object image can be {(suitable style: warm), (product color tendency: orange), (suitable background type: hand-drawn flowers)…}, and then the structured description text is converted into a corresponding vector, etc.
[0058] In this way, the similarity between the subject object image and each element can be calculated by calculating the distance between the vector corresponding to the subject object image and the vector corresponding to each element in the element asset library, and then the elements respectively matched with the subject object image under multiple types can be determined. After that, the subject object image and the matched elements can be organized according to the preset picture description protocol to generate a target picture.
[0059] In this way, mainly through the AI large model, the subject object image input by the user and various elements in the asset library are labeled, and then converted into mathematical expression vectors, and then the matching elements are calculated through distance calculation between vectors, and finally the subject object image and various types of matching elements are organized into specific target pictures. This ensures the quality of the generated image. Of course, in order to further improve the quality, the AI large model can also be controlled to label the subject object image or elements within a certain range, that is, the labels used for labeling are not randomly generated, but selected from multiple given labels by the AI, and so on. Through the above method, it is beneficial to improve the professionalism of the specific target picture, and thus improve the probability of its direct application in professional scenarios such as marketing material delivery.
[0060] In addition, in order to further improve the quality or professionalism of the generated target picture, the rendering method in the specific picture generation process is improved in the embodiments of the present disclosure. In the traditional picture design system, it is usually described in the form of layers and the coordinates of the specific content in the canvas are defined, and so on. However, due to the various sizes of subject object images in different user requirements, the sizes of the required display text, components, etc. are also different, so it may be difficult to guarantee the aesthetics and harmony of the picture in the description method according to the absolute coordinates.
[0061] In view of the above situation, the embodiments of the present disclosure define the specific layout scheme class element as a container and define the alignment of the specific content class element in the container, that is, in the embodiments of the present disclosure, the specific content class element is not placed in the absolute position of the canvas according to the coordinate information, but can be filled into the container according to the pre-defined alignment. This method is a structured relative positioning method, and can also realize the decoupling of layout, content and style, so as to improve the aesthetics of the specific generated picture, the harmony of the overall picture, and further improve the professionalism. For more specific implementation methods, detailed descriptions will be given in the following.
[0062] From the perspective of system architecture, as shown in FIG. 1, the embodiment of the present disclosure can first include an element asset library, which can include layouts, backgrounds, fonts, components, texts, color matching and other elements of different types, each type can include multiple different element instances. In specific implementation, a layout editor and a responsive component editor can also be provided to design specific layouts and component class elements and save them to the element asset library. In addition, a matcher can also be included, which can use the capabilities of AI large models to perform multi-modal (including image, text and other modal content) semantic tagging on the subject object image submitted by the user and each element in the element asset library, and convert it into a mathematical expression vector (Embedding). Then, the distance between vectors can be calculated to determine the matching elements of multiple types with the subject object image. Among them, the color matching elements can be generated programmatically after the matching of other elements is completed. Then, the renderer can organize the subject object image and the matching elements according to the preset picture description protocol to generate a specific target picture.
[0063] The specific implementation scheme provided by the embodiment of the present disclosure will be described in detail below.
[0064] Embodiment one
[0065] First, the present disclosure provides a method for generating a picture, as shown in FIG. 2, which can include:
[0066] S201: Determine the user's picture generation requirement information, the generation requirement information at least includes a subject object image.
[0067] The scheme provided by the embodiment of the present disclosure can be applied to professional picture generation scenarios. Specifically, the user can include marketing personnel in the system, and can also include merchants and designer personnel. Of course, in actual application, the scheme can also be applied to ordinary life scenarios, etc., at this time, the specific user can also be an ordinary consumer user, etc.
[0068] An operation portal for initiating a picture generation request can be provided to the user, and the user can initiate a specific generation request through this operation portal and submit specific generation requirement information. In the scenario of the present disclosure, the specific generation requirement information at least includes a subject object image. Among them, the subject object image can have different definitions in different scenarios. For example, in the scenario of generating marketing material pictures, the specific subject object image can be a product image, and in other scenarios, the specific subject object image can also be a person image, an object image, an animal image, etc.
[0069] In the marketing material related picture generation scenario, the user can directly upload a certain product picture when initiating the request, or if it is a product that has been published in the current product information service system, the user can also specify the ID or product detail page URL of the product, etc., and the system selects a certain product picture from the product text details for picture generation, etc.
[0070] In addition, in an optional manner, the specific generation requirement information can also include copy information, which can include information such as the main value point information of the product that needs to be reflected in the target picture, etc. Alternatively, if the user has a more specific requirement, the specific requirement can also be described in more detail through specific copy, for example, the specific scene, style, color matching, etc. required can be described. Alternatively, in the marketing material delivery scenario, the specific delivery channel information can also be described through the copy, since some common delivery channels usually have official standards for the size of the delivered marketing material picture, etc., therefore, in the case of knowing this information, the target picture generated can also be more in line with the specific delivery requirement, which is also one of the professional manifestations of the generated picture.
[0071] S202: generating semantic labels for the subject object image through an artificial intelligence (AI) large model, so as to obtain a description text for describing the features of the subject object image according to the semantic labels, and convert the description text into a first vector.
[0072] After receiving the requirement information of the user, the AI large model can be used to generate semantic labels for the subject object image in the requirement information, so as to obtain a description text for describing the features of the subject object image. Then, the description text can be converted into a first vector expressed in mathematics.
[0073] In a specific implementation, in order to improve the subsequent matching accuracy, the labels output by the AI large model can be controlled, for example, a plurality of label instances in multiple dimensions can be enumerated in advance, and a prompt text for interacting with the AI large model can be constructed. Then, the subject object image of the user's requirement and the above prompt text can be input into the AI large model for labeling by the AI large model.
[0074] For example, in the scenario of generating a marketing material picture, the specific subject object image is a product picture of a certain product, and when labeling the product picture by the AI large model, the prompt text input into the AI large model can include the following content:
[0075] "Show product category: what is the product category in the current picture? You can refer to but not limited to: notebook computer, hair dryer, eye shadow, ……
[0076] Background style suitable for the product: What is the current product adapted to the background style, which can be but not limited to: calm, warm, minimalist, complex, fashionable, hot……;
[0077] Background type suitable for the product: What is the current product adapted to the background type, which can be selected from the following but not limited to: white, black, gray, clean single color, single color background, background with gradient color, clean gradient color, deep gradient color, vector geometry, vector geometry color block segmentation, vector curve color block segmentation, curve line, hand-drawn plants, hand-drawn flowers, hand-drawn animals, hand-drawn scenes, light lines, grid space, pure color light and shadow, plant shadow, blue sky and clouds, 3D geometry, 3D flowers, 3D sphere, 3D rendering abstract graphics……;
[0078] Background content suitable for the product: Please use your imagination, what kind of background is suitable for the current product, give a detailed text description;
[0079] Product color system tendency: What is the color system of the product in the current product picture, which can be referred to but not limited to: yellow-green color system, black and white color system, blue-green color system, pink-purple color system;
[0080] Color sorting: What are the three colors with the largest proportion in the current product picture, give an array of 3 main colors, and sort them according to the proportion of the 3 colors in the product picture from large to small;
[0081] Recommended background: If the current product picture is combined with other products of the same type to form a poster design, what kind of background should the product be matched with, please give a specific description text;
[0082] Holiday suitable for the product: What holiday or promotion node marketing is the current product adapted to, please give a specific text description, if not please give ‘ / ’ directly;
[0083] Product target audience: What is the current product suitable for the target audience, and what are the preferences of the target audience for the product background, please give a detailed text description.
[0084] Through the above prompt text, the AI large model can generate more standardized labels for specific subject object images. Of course, if none of the listed label examples are suitable, the AI large model can also generate other labels.
[0085] S203: Determine the elements of multiple types respectively matched with the subject object image according to the first vector corresponding to the subject object image and the second vectors respectively corresponding to multiple different types of elements in the pre-established element asset library; wherein, the elements are the elements required in the generation of pictures, and the second vectors are obtained by converting the description text used to describe the features of the elements after generating the semantic labels of the elements by the AI large model and obtaining the description text according to the semantic labels.
[0086] After obtaining the first vector corresponding to the subject object image, distance calculation can be performed with the second vectors respectively corresponding to multiple different types of elements in the pre-established element asset library. As described above, the specific elements can be multiple types of elements required in the generation of pictures, including layout, background, font, component, copy, color matching, etc. The layout is used to define the skeleton structure of the picture arrangement style; the background is the picture with the largest proportion in the picture and the bottom layer; the component refers to the relatively secondary level copy in the picture except the title, including operation buttons, such as promotion, logistics, discount, after-sales, action point, selling point function, etc.
[0087] For the above various types of elements, the AI large model can also be used to generate semantic labels for the elements to obtain description texts used to describe the features of the elements, and then the corresponding second vectors in mathematical expressions can be obtained by converting the description texts. That is, each element instance in the element asset library can also correspond to its own semantic label and the second vector obtained by conversion. Specifically, a management entrance of the element asset library can be provided in the background, and the designers or business users in the system can add new element instances to the element asset library, and each time a new element instance is added, the AI large model can add a label for it and convert it into a corresponding vector. The specific way of adding labels will be described in detail later.
[0088] In the above, in order to further improve the professional degree of the generated picture, the design and expression manner of the layout class element are further provided in the embodiments of the present disclosure. Specifically, first, various types of elements in the element asset library can be classified. Specifically, since the layout is usually used to define the canvas layout style skeleton structure, the components, the text and the user input subject image usually belong to the foreground content that needs to be filled into the canvas, and the font, the color matching and the like are mainly used to express the style of the specific content. Therefore, the above elements can be classified into three categories, i.e., the layout, the content and the style. The content class element can further include multiple subtypes, for example, can include the image class, the text class and the component class, etc. The style class element can specifically include the font, the color matching and the like. In addition, since the background can usually play a role in setting the tone of the whole picture, the background can be determined first, and then the font, the color matching of the component and the like can be determined. Therefore, the background can also be classified as a separate element. In this way, various types of elements in the element asset library can be classified into four categories, i.e., the layout, the content, the style and the background. In the embodiments of the present disclosure, the structured relative positioning design of the layout element can be used to realize the mutual decoupling between the above elements, so that each part can be dynamically adjustable, and at the same time, it can be ensured that the generated picture can be guaranteed in terms of the aesthetic degree, the picture harmony and the like when each part is changed.
[0089] In the embodiments of the present disclosure, the structured relative positioning design of the layout refers to that the relative position of the content class element in the canvas can be defined by the layout element. In this way, when the renderer organizes various matched elements subsequently, the matched content class elements can be filled into the corresponding position in the canvas according to the definition of the matched target layout scheme class element.
[0090] Specifically, the layout element can be used to divide the canvas into multiple information regions, and define the subtypes of the content class elements needed to be filled in each information region and the alignment manner information of the content class elements in the information region, so as to fill the content class elements of the corresponding subtype into the corresponding information region according to the alignment manner. That is to say, in the embodiments of the present disclosure, the specific content class element is filled into a certain information region in the specific canvas according to the pre-defined alignment manner, for example, is centered aligned, or is aligned to the left or the right, etc. In this way, no matter how the size of the specific content class element changes, the same alignment manner can be maintained in the corresponding information region, so as to guarantee the overall picture harmony and the like.
[0091] For example, in one way, the canvas can be divided into 3 information regions: a main information area, a secondary information area, and a main content area, a total of three parts, i.e., the screen is divided into 3 parts, so that a total of 6 main formats can be enumerated, which are the basic formats of three-segment structured positioning. Specifically, as shown in FIG. 3. Of course, FIG. 3 only exemplarily shows the division manner of different information regions, and in specific implementation, the size, proportion, etc. of the specific canvas can also be adjusted, for example, each can be respectively lengthened or shortened by a certain proportion in the horizontal or vertical direction, etc.
[0092] In an optional implementation, on the basis of the above structural design paradigm based on the design layout basic specification and principle, the flex mode can also be integrated to realize the definition of the format structure by defining the main and auxiliary axes and forced alignment, aiming to realize the deep integration with the proximity principle, similarity principle and continuity principle in the design gestalt theory. That is, the alignment manner information of the content class element in the information region can include: the alignment manner of the content class element in the main axis direction and / or the auxiliary axis direction of the information region, wherein the main axis direction is the arrangement direction of the content class element in the information region, and the auxiliary axis direction is the vertical direction of the main axis direction.
[0093] In the above manner, each information region in the layout can be controlled by the following parameters:
[0094] 1. axis: main axis direction, i.e., the arrangement direction of the content class element in the information region, for example, assuming that three pieces of text content need to be displayed in an information region, whether the three pieces of text content are arranged horizontally or vertically can be determined by the parameter. Specifically, the value of the parameter can be "x" or "y", "x" indicating that the main axis direction of the region is horizontal, and "y" indicating that the main axis direction of the region is vertical.
[0095] 2. justify: the main axis alignment manner of the information region, the value of the parameter can include: "center", "start", "end", "between", "around".
[0096] wherein "center" is the center alignment in the main axis direction, assuming the current main axis direction is horizontal, when the alignment is "center", the display effect of the content class element in the information area can be as shown in FIG. 4(A). Among them, the dashed box of the rectangle represents an information area in the layout, and three pieces of content, such as "One", "Two", and "Three has extra text", need to be displayed in the information area. Of course, in the example shown in FIG. 4(A), the third piece of content will wrap at the end of each word, which is controlled by the content element. When displaying the above three pieces of content in a certain information area of the canvas, the specific position of the three pieces of content in the information area can be determined by the parameter values of the parameters axis, justify and the like. When the parameter value of axis is "x" and the parameter value of justify is "center", the display effect of the above three pieces of content in the information area can be as shown in FIG. 4(A). That is, the three pieces of content are arranged from left to right in the horizontal direction, and the three pieces of content are displayed in the center of the area.
[0097] "start" is the head alignment in the main axis direction, assuming the current main axis direction is horizontal, when the alignment is "start", the content class element is aligned to the left in the information area, for example, the specific effect can be as shown in FIG. 4(B).
[0098] "end" is the tail alignment in the main axis direction, assuming the current main axis direction is horizontal, when the alignment is "end", the content class element is aligned to the right in the information area, for example, the specific effect can be as shown in FIG. 4(C).
[0099] "between" is the head-tail alignment in the main axis direction, assuming the current main axis direction is horizontal, when the alignment is "between", the content class element will fill the entire area in the information area, for example, the specific effect can be as shown in FIG. 4(D).
[0100] "around" is the average distribution in the main axis direction, assuming the current main axis direction is horizontal, when the alignment is "around", the content class element will be evenly distributed in the information area, for example, the specific effect can be as shown in FIG. 4(E).
[0101] 3. align: the auxiliary axis alignment mode of the information area. That is, the specific content class element can be aligned in the main axis direction of the information area, and can also be aligned in the auxiliary axis direction in a certain way. Among them, the auxiliary axis is the vertical direction of the main axis, for example, if the main axis is the horizontal direction, that is, the x direction, the vertical direction, that is, the y direction, is the auxiliary axis, and so on. The parameter value of the auxiliary axis alignment mode can also include "center", "start", "end", "between", "around", which respectively represent center alignment, head alignment, tail alignment, head-tail alignment, and average distribution.
[0102] For example, assuming that the current main axis direction is horizontal, the main axis alignment direction is "center", that is, center alignment, and the alignment mode in the auxiliary axis direction is "center", "start", "end", and "between", respectively. The corresponding display effect can be as shown in Figures 5(A), (B), (C), and (D), respectively.
[0103] In addition, the parameters of the specific information area can also include:
[0104] 4. "pl": the parameter indicates how much area is left in the left part of the current area;
[0105] 5. "pt": the parameter indicates how much area is left in the top part of the current area;
[0106] 6. "pb": the parameter indicates how much area is left in the bottom part of the current area;
[0107] 7. "pr": the parameter indicates how much area is left in the right part of the current area;
[0108] 8. "gap": the parameter indicates the spacing of the elements in the current area;
[0109] 9. "percent": the parameter indicates the percentage of the current area in the total area;
[0110] 10. "background": the parameter indicates whether the current area needs to be filled with a background color.
[0111] The above is the definition of the parameters in each information area. The content part can include product images, text (including titles, value points, brand logos, etc.), components, and other parts. In addition, it can also include background elements. Each element can be freely placed in each information area, and the information area defined by the above parameters belongs to a container nature. The specific content class element in the information area can be automatically aligned by the above alignment mode rules.
[0112] In this way, a simple rule set can be taken as the core, which is not only suitable for typesetting needs in a wide range of marketing material generation scenarios, but also can effectively reduce costs, realize the rapid generation of intelligent marketing materials, and has strong expansibility, and can be seamlessly connected with various AI backlink technologies (for example, Stable Diffusion image generation technology, large language models, etc.). In addition, since the mathematical coordinate calculation is removed, only the placement position and alignment method of the element need to be considered during layout, thus simplifying the design logic of absolute layout typesetting, which is more suitable for logical operation scenarios of language models. Moreover, this strict alignment method is more in line with rigorous design specifications and design aesthetics, in addition, through structured layout, elements can naturally adapt to the size of the canvas dynamically, and the size adjustment is flexible.
[0113] The above mentioned canvas size dynamic adaptation is expanded here. In the embodiments of the present disclosure, a plurality of specific elements can be pre-stored in the element asset library, which includes layout type elements. The definition of the layout type elements is as described above. Each specific layout type element instance has its own size, aspect ratio, etc. For example, the aspect ratio of a certain layout type element instance is 1:1, the aspect ratio of another layout type element instance is 3:4, and the aspect ratio of another layout type element instance is 16:9, etc. In principle, when generating an image based on user needs, if a layout element with a certain aspect ratio is needed, an element with the same aspect ratio can be selected for generation, which is the most ideal state. However, in actual application, different user needs correspond to various generated image aspect ratios and sizes, some of which may be uncommon sizes, such as 3.5:4, etc. This makes it difficult to exhaustively list all possible user required sizes and design corresponding layout type element instances. On the other hand, since the structured layout design method described above is adopted for layout type elements in the embodiments of the present disclosure, this element can naturally adapt to the size of the canvas dynamically, that is, as long as the actual required size deviates too much from the size of a certain layout type element, the generation of a specific image can be based on the parameters of the layout type element, and the position of the content element in the generated image can still be maintained within a reasonable range. Of course, if the actual required size deviates too much from the size of a certain layout type element, it is no longer suitable for image generation through that layout type element, at which point another layout type element with a size closer to the actual required size can be selected for image generation.
[0114] Therefore, based on the above, in the specific implementation, a plurality of breakpoints can be set in advance in the size and / or aspect ratio dimension, and the corresponding layout scheme element is configured in the element asset library for the plurality of breakpoints. For example, in the aspect ratio dimension, the specific breakpoint position can include 1:1, 3:4, 16:9, etc., and of course, some more extreme cases can also be set, such as a very narrow horizontal or vertical long strip layout element, etc. Other non-breakpoint positions can not necessarily be configured with corresponding layout elements in the element asset library. Specifically, when generating a picture according to the generation requirement information, the target size and / or aspect ratio of the picture to be generated can be determined first. If the target size and / or aspect ratio does not hit the breakpoint, the layout scheme element with the closest size and / or aspect ratio can be selected, and the position of the content element is adaptively determined according to the structured relative positioning design information defined by the layout scheme element. It can be seen that through the scheme provided by the embodiments of the present disclosure, intelligent dynamic adjustment of multiple different sizes can be realized, and consistent visual experience can be provided for users in diversified devices and scenes.
[0115] Among them, the size and / or aspect ratio of the picture to be generated in the specific generation requirement can be determined in various ways. For example, in one way, the user can specify the specific size and / or aspect ratio when inputting the requirement information, that is, the text requirement information input can include text description content about the specific size and / or aspect ratio. Alternatively, in another way, for the marketing material delivery scene, the user can also specify the specific delivery channel in the input requirement. Since the specific delivery channel usually has standard official requirements for the size and / or aspect ratio of the specific marketing material, these standard information can be collected by querying in the public network, and therefore, the size and / or aspect ratio of the picture to be generated by the user can also be determined according to the user-specified delivery channel information.
[0116] The above describes the definition of the layout type element, and in addition, the component class element is also improved in the embodiment of the disclosure. In the traditional component design, the background frame of the component is usually drawn first (for example, the component is a button, and the button needs a background frame of a round rectangle, and some text needs to be displayed on the button to express the function of the button, such as "go and see" action point text), and then the text is filled in the background frame. However, the following situation often occurs: if the text to be filled is relatively long, it may exceed the background frame, so that the text cannot be fully displayed, or the beauty of the whole picture is affected. In view of this situation, the scheme provided by the embodiment of the disclosure is to first determine the size of the background frame required to carry the text content and the font size information of the text content in the component according to the text content and the font size information of the text content in the component, and then draw the background frame of the component, so that the size of the background frame of the component can be adaptively changed with the text displayed in the component.
[0117] In addition, the style of the responsive component can also be specified in the form of parameters. For example, the style of the component can be controlled by the following variables: serial number, stroke line type, outer stroke line type, stroke thickness, outer stroke thickness, corner radius, projection position, projection color, stroke color, outer stroke color, font color, font component serial number, background color, main color tendency, and the like. This implementation of specifying the style of the responsive component in the form of parameters can better combine the structural positioning paradigm operation advantages, aesthetic advantages, and expansion advantages of the layout type element, and the style can be more flexible to adapt to the canvas ratio and the number of text.
[0118] For the background, a background picture library can be preset according to multiple dimensions such as country, industry, holiday node, and delivery platform, to better adapt to the customized delivery needs of different countries, crowds, platforms, and holidays.
[0119] The above describes the design principles of various types of elements. Based on the above principles, various specific element instances can be added to the element asset library, for example, a specific layout type element instance can be added, or a specific background picture can be added, or a screenshot of a specific font can be added, a specific color matching scheme can be added, and the like.
[0120] After adding specific element instances to the element asset library, an AI large model can be used to generate semantic tags for the element instances, to obtain structured description texts for describing the features of the elements.
[0121] In a specific implementation, a specific element instance can be input into the AI large model to enable the AI large model to add semantic labels to the specific element. To enable the labels added by the AI large model to better serve subsequent vector distance calculation, some exemplary labels can also be provided to the AI large model, which can be written in the prompt text to control the output of the AI large model. For example, for elements of the layout class, when adding a certain element instance in the asset library, in addition to the structured information of the specific element, including the division method of each information area, a picture example based on the layout element can also be provided. At this time, the picture example and the above constructed prompt text can be input to the AI large model to add specific semantic labels to the element instance. For example, in a specific implementation, the prompt text constructed for the layout element can be:
[0122] "Scene specification: Which marketing scenarios is the current layout suitable for, such as marketing channels A, B, C, ……
[0123] Target audience: What is the current layout suitable for, and what are the preferences of the target audience for the product background? Please provide a detailed text description;
[0124] Style: Which of the following is the current layout suitable for, highlighting the product, highlighting the discount information, highlighting the title information, or highlighting the use scenario? Please provide a detailed text description;
[0125] Product aspect ratio: What aspect ratio of product does the current layout apply to, such as square, long and horizontal products;
[0126] Number of products: How many products does the current layout apply to in marketing pictures."
[0127] In this way, the AI large model can give corresponding labels in the above scene specification, target audience, style, product aspect ratio, and number of products dimensions.
[0128] Similarly, for elements of the background class, a specific background element instance is usually a specific background picture, so the background picture and the pre-designed prompt text can be input to the AI large model to label the background instance. For example, in the scenario of generating marketing material pictures, the prompt text constructed for the background element can be:
[0129] "Product category (productsType): Give a product category: What category does the displayed product belong to, which can be referred to but not limited to: notebook computer, hair dryer, eye shadow, ……
[0130] backgroundStyle: What is the visual style of the current marketing material background? It can be but not limited to: calm, warm, minimalist, complex, fashionable, hot
[0131] backgroundType: What type is the current marketing material background? It can be but not limited to: white, black, gray, clean monochrome, monochrome undercoat, undercoat with gradient color, clean gradient color, deep gradient color, vector geometric figure, vector geometric color block segmentation, vector curve color block segmentation, curve line, hand-drawn plants, hand-drawn flowers, hand-drawn animals, hand-drawn scenes, light lines, grid space, pure color light and shadow, plant shadow, blue sky and clouds…
[0132] backgroundContent: What is the visual style of the current marketing material background? Please give a detailed text description;
[0133] backgroundColorTendency: What is the color system of the current marketing material? It can be but not limited to: yellow-green color system, black and white color system, blue-green color system, pink-purple color system;
[0134] backgroundColorSorting: What are the three colors with the highest proportion in the current marketing material background? Give an array of 3 main colors in HEX, and sort them according to the proportion in the marketing material from large to small;
[0135] matchingProducts: If the background of the current marketing material is used for poster design, what products should it match? Please give specific options text;
[0136] festivalMatching: What festival or promotion node does the background of the current marketing material adapt to? Please give specific options text, or directly give " / " if there is none;
[0137] user: If the background of the current marketing material is used to design other marketing materials, what kind of user profile is suitable for designing marketing materials with this background? Please give a detailed text description of the visual preference characteristics of this user profile.
[0138] The AI large model can give corresponding labels for specific background elements in the dimensions of the above-mentioned display product categories, background types, background contents, background color system tendencies, color sorting, matching products, festival matching, and user groups.
[0139] For elements of style, font class, generally used to describe the style, font of text content such as title, or the style, font of component, etc., therefore, in the embodiment of the disclosure, in order to be able to mark the specific style, font, generally can be expressed by the text content, component screenshot to which the style, font is applied. In this way, the specific screenshot can not only express the font, style, but also reflect the specific application of the text content, component, that is, it can be known that the font, style is used for what text content, component, so as to facilitate the marking of the specific font, style class element. After uploading the specific screenshot, the screenshot and the prompt text can be input into the AI large model, so as to mark the style, font class element. For example, also in the scene of generating marketing material picture, the prompt text constructed for the style, font class element of the title text content can be:
[0140] "Main and sub-title content: what is the main and sub-title information content of the current marketing picture, please give the specific text information, if not please give " / " directly;
[0141] Picture position of main and sub-title: where is the position of the main and sub-title of the current marketing picture in the picture, which can be but not limited to the following options: picture center, picture left side, picture right side, picture upper left, picture upper right, if not please give " / " directly;
[0142] Main and sub-title font: what is the font of the main and sub-title of the current marketing picture, select from the following options: bold, bold italic, serif, handwriting, if not please give " / " directly;
[0143] Main and sub-title font style: what is the font style of the main and sub-title of the current marketing picture, select from the following options: with stroke, with projection, with stroke and projection, if not please give " / " directly;
[0144] Main and sub-title font color: what is the font color of the main and sub-title of the current marketing picture, please give the specific array of HEX, if there is no main title in the picture please give " / " directly;
[0145] Target audience of title content: what is the target audience of the main and sub-title content of the current marketing picture, please describe a text specifically."
[0146] The prompt text designed for the style, font class element of the component can include:
[0147] "Price (price): what is the price of the product given in the current marketing material, please give the specific text, if not please fill in " / "
[0148] Price font (priceFont): What is the font of the displayed price in the current marketing material? Please choose one from the following options: bold, bold italic, serif, handwriting, and if there is none, write ' / '
[0149] Price font color (priceFontColor): What is the font color of the displayed price in the current marketing material? Please give the specific HEX value
[0150] Price font style (priceFontStyle): What is the font style of the displayed price in the current marketing material? Please choose one from the following options: with stroke, with shadow, with stroke and shadow, and if there is none, fill in ' / '
[0151] Price block color (priceFontBlockColor): What is the color of the color block under the displayed price font in the current marketing material? Please give the specific HEX value, if the price font is under the background, please directly give ' / '
[0152] Price block shape (priceFontBlockShape): What is the shape of the color block under the displayed price font in the current marketing material? Please give a specific shape description, if the price font is under the background, please directly give ' / '
[0153] Price block shape style (priceFontBlockStyle): What is the shape style of the color block under the displayed price font in the current marketing material? Please choose one from the following options: single block, gradient color, single block with stroke, single block with shadow, single block with stroke and shadow, gradient color with stroke, gradient color with shadow, gradient color with stroke and shadow, and if there is none, fill in ' / '
[0154] Price audience (priceFontUser): With price content, price block, price block shape, and price block shape style as price label, what is the portrait of the current price label audience? Please give a specific description text
[0155] Price label screen ratio (priceScreenRatio): With price content, price block, price block shape, and price block shape style as price label, what is the ratio of the current price label in the screen? You can refer to but not limited to: 10%, 20%;
[0156] …
[0157] Regarding the style, font class elements of the components, in addition to the above-mentioned dimensions, other dimensions can also be included when marking, such as price tag picture position, discount, discount font, discount font color, discount font style, discount color block color, discount color block shape, discount color block shape style, discount audience group, discount label picture proportion, discount label picture position, product selling point, product selling point font, product selling point font color, and the like, which will not be listed one by one here.
[0158] In summary, regarding the layout, background, style, font, and other types of elements, they can be marked respectively by the above-mentioned method, and then the obtained description text can be converted into a second vector expressed in mathematics for calculation with the first vector corresponding to the image of the subject object input by the user. For example, in a specific implementation, the distance between the first vector corresponding to the image of the subject object (which can be referred to as Embedding1) and the second vector corresponding to each type of element (which can be referred to as Embedding2) can be calculated according to the cosine similarity, so as to obtain the similarity calculation result between the two and sort them, and then obtain the matching result closest in semantics. The matching process can be as shown in FIG. 6.
[0159] This way of element matching by AI large model semantic marking of subject object image and various design elements, and then converting to mathematical expression vector for distance calculation, is to apply the semantic analysis of Kansei engineering in the field of intelligent layout of picture generation, and can use AI large model to make perceptual description of different types of elements based on the theory of Kansei engineering of semantic vector, and combine objective parameters to realize intelligent matching of visual elements of different styles, industry categories, and target audience preferences. This method produces intelligent visual matching organization logic, which not only improves the aesthetic value of design practice, but also significantly improves the visual appeal and conversion rate of pictures to users (design accumulation + business data), and provides strong support for the successful implementation of marketing strategies.
[0160] In addition, regarding the color matching elements, the matching process can be completed after the matching of background and other elements is completed using programmatic matching rules. Specifically, an automatic color matching logic based on the background main color as the starting point and the HSL (Hue Saturation Lightness, hue, saturation, and lightness, which is a method of representing points in the RGB color model in cylindrical coordinates) color value rule can be used. For example, the specific logic can be as follows:
[0161] First, color plate definition can be performed, which can specifically include:
[0162] primary: main color
[0163] primary Contanier: primary light color
[0164] primary Dark: primary dark color
[0165] similar: similar color
[0166] similar Contanier: similar light color
[0167] similar Dark: similar dark color
[0168] contrast: contrast color
[0169] contrast Contanier: contrast light color
[0170] contrast Dark: contrast dark color
[0171] black: black
[0172] white: white
[0173] Then, initialize primary color: assume the primary color of the input background element (i.e., the background element that matches the current input subject object image successfully) is mainColor;
[0174] Finally, generate primary color system: define a set of HSL offsets for each color value in the color palette, which can cause mainColor to change in H (hue), S (saturation), and L (lightness) by a specified amount, so that each color in the color palette meets the defined requirements.
[0175] For example, in one example, the specific code mainly generates a series of color variations based on a given color (color), including primary color, contrast color, and similar color, as well as their corresponding container color and dark color. For example, it can include:
[0176] Generate primary color system:
[0177] primary: set the saturation (s) of mainColor to 0.9 and the brightness (l) to 0.6 to get a high-saturation and medium-brightness primary hue.
[0178] primaryContainer: increase the brightness of primary color by 3 units to create a container or background color that looks brighter and contrasts with the primary color.
[0179] primaryDark: reduce the brightness of primary color by 2.5 units to create a darker version for dark themes or places that need emphasis.
[0180] Generate similar color scheme:
[0181] similar: subtract 40 from the hue (h) of mainColor, and set the saturation and lightness to 0.9 and 0.6 respectively, to create a color similar to the main color but with differences.
[0182] The generation logic of similarContainer and similarDark is the same as the main color scheme, but based on the similar color.
[0183] Generate contrast color scheme:
[0184] contrast: increase the hue of mainColor by 180 degrees, so that the resulting color is opposite to the original color on the color wheel, forming a sharp contrast. The saturation and lightness are also set to 0.9 and 0.6. The generation logic of contrastContainer and contrastDark is the same as the previous two, but based on the contrast color.
[0185] S204: Organize the subject object image and the matched elements according to the preset picture description protocol to generate a target picture.
[0186] After matching multiple types of elements, the subject object image and the matched elements can be organized according to the preset picture description protocol to generate a target picture. That is, the user inputs the required subject object image, and the intelligent picture generation engine automatically generates the picture by calling the layout, background, text, font, component, color matching, and other types of design elements that match the subject object image. Among them, in the way of using the aforementioned structured design layout element, the picture description protocol can be composed of five sub-modules, which are:
[0187] a. Layout module: mainly used to describe the basic layout framework.
[0188] b. Content module: mainly used to describe the specific content filled between the information areas of the canvas.
[0189] c. Background module: mainly used to describe background information.
[0190] d. Style module: mainly used to describe the configuration information of each responsive component.
[0191] e. Color matching module: mainly used to record the generated color matching table.
[0192] The subject object image input by the user, the matched elements of various types, and the picture description protocol can all be provided to the renderer, which organizes the subject object image and the matched elements of various types according to the picture description protocol to generate a final picture for output. For example, the generated picture can include a marketing material picture, which can be provided to an operator for delivery to a specific delivery channel, or can be a design poster, and the like.
[0193] In summary, according to the embodiments of the present disclosure, an element asset library can be established in advance for the user's demand for generating a picture, which can include various types of design elements. For each specific element instance, an AI large model can be used to generate a semantic label for the element to obtain description text for describing the features of the element, which can then be converted into a second vector for mathematical expression. After the user inputs a specific generation requirement, an AI large model can also be used to generate a semantic label for the subject object image specified by the user to obtain description text for describing the features of the subject object image, and the description text can be converted into a first vector. In this way, the elements of various types that are respectively matched with the subject object image can be determined according to the first vector corresponding to the subject object image and the second vectors respectively corresponding to the elements of various types. Then, the subject object image and the matched elements can be organized according to a preset picture description protocol to generate a target picture. In this way, not only is the automatic generation of pictures achieved, but the semantic analysis of ergonomics can also be applied to the picture intelligent layout scene. The method produces intelligent visual matching and organization logic, which can ensure that the automatically generated picture has an aesthetic value and professionalism in design practice, thereby helping to improve the visual appeal of the picture to the user and even improve conversion rates and other indicators, providing strong support for the successful implementation of marketing strategies, and the like.
[0194] In the preferred implementation, on the basis of the semantic labeling and kansei engineering element matching of the multi-modal language model to realize the automation of the production of design materials, it can also be combined with the flexible layout (Flex) flexible layout structure, responsive design components, etc., and the design rule set based on the summary of the design method can be realized, and the AI large model logical thinking ability is interspersed in it, which is beneficial to realize intelligent dynamic adjustment under various sizes, and ensure that users can provide consistent visual experience in diversified devices and scenes. For example, assuming that the user input subject object image is a picture about a teddy bear toy, under the condition that the size of the picture is different, the size / aspect ratio of the generated picture is different, etc., the generated picture can be as shown in FIG. 7, which gives the picture generation results of 7 different sizes / aspect ratios. Although the positions of the subject object image, the text, the components, etc. in each picture, the arrangement mode between different contents, the background picture, the font, the color matching, etc. are all different, but in terms of picture aesthetics, overall coordination, etc. are reasonable, and belong to professional level picture generation results.
[0195] Embodiment two
[0196] This embodiment two is corresponding to embodiment one, from the perspective of the client, a method for generating a picture is provided, as shown in FIG. 8, which can include:
[0197] S801: receiving user generation demand information of a picture, the generation demand information at least including a subject object image;
[0198] S802: submitting the generation demand information to a server, the server being configured to generate a semantic label for the subject object image by an AI large model, so as to obtain a description text for describing the features of the subject object image according to the semantic label, and then convert the description text into a first vector; according to the first vector corresponding to the subject object image and the second vectors corresponding to a plurality of different types of elements in a pre-established element asset library, respectively determining a plurality of types of elements respectively matched with the subject object image, and organizing the subject object image and the matched elements according to a preset picture description protocol to generate a target picture; wherein the element is an element required in the generation process of the picture, and the second vector is obtained by converting the description text after generating a semantic label for the element by an AI large model, and obtaining a description text for describing the features of the element according to the semantic label;
[0199] S803: receiving and displaying the generated target picture.
[0200] It should be noted that the embodiments of the present disclosure can involve the use of user data. In actual application, user-specific personal data can be used in the schemes described herein in a manner that complies with the applicable legal requirements of the country (for example, with the explicit consent of the user, with the actual notification to the user, etc.) and within the scope permitted by the applicable laws and regulations.
[0201] Corresponding to the foregoing method embodiment I, the embodiments of the present disclosure also provide a device for generating a picture, which can include:
[0202] a demand determination unit configured to determine user generation demand information of a picture, the generation demand information including at least a subject object image;
[0203] a first vector generation unit configured to generate a semantic label for the subject object image by using an artificial intelligence (AI) large model, so as to obtain a description text for describing features of the subject object image according to the semantic label, and convert the description text into a first vector;
[0204] a vector calculation unit configured to determine elements of different types respectively matched with the subject object image according to the first vector corresponding to the subject object image and second vectors respectively corresponding to a plurality of different types of elements in a pre-established element asset library; wherein the elements are elements required in the process of generating a picture, and the second vectors are obtained by converting a description text for describing features of the elements after generating a semantic label for the elements by using the AI large model and obtaining the description text according to the semantic label;
[0205] a picture generation unit configured to organize the subject object image and the matched elements according to a preset picture description protocol to generate a target picture.
[0206] In the process of generating a semantic label for the subject object image, the generated semantic label includes a category of the subject object, a color tendency, a suitable occasion, a style, and / or a crowd.
[0207] In the process of generating a semantic label for the elements, the generated semantic label includes a suitable scene, a style, a crowd, a category of the subject object, an aspect ratio, and / or a quantity of the elements.
[0208] Specifically, the elements in the element asset library include layout scheme elements, content elements, style elements, and background elements, wherein the content elements include a plurality of subtypes, the subtypes include image types, text types, or component types, the style elements are used to describe style configuration information of text and / or components, and the style includes a font and / or a color matching;
[0209] The layout scheme class element is used to define a canvas layout format skeleton structure, and the relative positions of the content class elements in the canvas are defined through a structured relative positioning design method, so that the matched content class elements are filled into the corresponding positions in the canvas according to the definition of the matched target layout scheme class element.
[0210] Specifically, the layout scheme class element is used to divide the canvas into a plurality of information regions, and the subtypes of the content class elements required to be filled in each information region and the alignment manner information of the content class elements in the information region are defined respectively, so that the content class elements of the corresponding subtypes are filled into the corresponding information regions according to the alignment manner.
[0211] The alignment manner information of the content class elements in the information region includes:
[0212] The alignment manner of the content class elements in the main axis direction and / or the auxiliary axis direction of the information region, the main axis direction being the arrangement direction of the content class elements in the information region, and the auxiliary axis direction being the vertical direction of the main axis direction.
[0213] In a specific implementation, in the process of generating a picture according to the generated requirement information, for the content of the component class, the size of the background picture frame required to carry the text content and the font size information required to be displayed in the component are determined first, and then the background picture frame of the component is drawn, so that the size of the background picture frame of the component can be adaptively changed with the displayed text in the component.
[0214] In addition, a plurality of breakpoints can be set in advance in the size and / or aspect ratio dimension, and corresponding layout scheme class elements are configured in the element asset library for the plurality of breakpoints.
[0215] When generating a picture according to the generated requirement information, the target size and / or aspect ratio of the picture required to be generated can be determined, if the target size and / or aspect ratio does not hit the breakpoint, the layout scheme class element with the closest size and / or aspect ratio is selected, and the positions of the content class elements are adaptively determined according to the structured relative positioning design information defined by the layout scheme class element.
[0216] In the process of determining the matched element, the background element is determined, and the color matching element is determined through a programmatic generation method.
[0217] When the color matching element is determined through a programmatic generation method, the main color of the matched background class element is determined, and a group of offsets is defined from the main color as a starting point, the generated main color system is changed in hue, S saturation, and L lightness through the offsets, and the generated color version color meets the defined requirement.
[0218] In addition, the generated requirement information can further include copy content, the copy content including copy content to be added to the generated picture, and / or copy information describing a required scene, style, and required target population.
[0219] The subject object image includes a product image, and the generated target picture includes a marketing material picture for participating in a marketing activity.
[0220] Corresponding to Embodiment Two, the disclosure also provides a device for generating a picture, which can include:
[0221] A requirement receiving unit configured to receive picture generation requirement information of a user, the requirement information including at least a subject object image;
[0222] A requirement submitting unit configured to submit the requirement information to a server, the server being configured to generate semantic labels for the subject object image by using an AI large model, to convert description text describing features of the subject object image into a first vector based on the semantic labels, to determine elements of different types that match the subject object image under different types based on a second vector corresponding to each of the elements in an element asset library and the first vector corresponding to the subject object image, and to organize the subject object image and the matching elements according to a preset picture description protocol to generate a target picture. The elements are elements required in the picture generation process, and the second vector is obtained by converting description text describing features of the elements into a second vector based on semantic labels generated for the elements by using the AI large model.
[0223] A picture receiving unit configured to receive and display the generated target picture.
[0224] In addition, the disclosure also provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the method of any one of the preceding method embodiments.
[0225] An electronic device including:
[0226] one or more processors; and
[0227] a memory associated with the one or more processors, the memory being configured to store program instructions that, when executed by the one or more processors, perform the steps of the method of any one of the preceding method embodiments.
[0228] A computer program product comprising computer program / computer executable instructions to implement the steps of the method according to the preceding method embodiments when executed by a processor in an electronic device.
[0229] Fig. 9 shows an exemplary architecture of the electronic device, which can specifically include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920 can be communicatively connected through a communication bus 930.
[0230] The processor 910 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the present disclosure.
[0231] The memory 920 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device 900, a BIOS (Basic Input Output System) for controlling the low-level operation of the electronic device 900. In addition, a web browser 923, a data storage management system 924, and a generated image processing system 925, etc. can also be stored. The generated image processing system 925 can be an application program for implementing the above steps according to the embodiments of the present disclosure. In summary, when the technical solutions provided by the present disclosure are implemented by software or firmware, the relevant program codes are stored in the memory 920 and executed by the processor 910.
[0232] The input / output interface 913 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices can include a display, a speaker, a vibrator, an indicator light, etc.
[0233] The network interface 914 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through wired mode (such as USB, network cable, etc.), or can realize communication through wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0234] The bus 930 includes a path for transmitting information between various components (such as the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920) of the device.
[0235] It should be noted that although the above device only shows the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, the memory 920, the bus 930, etc., in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary to implement the present disclosure scheme, and does not have to contain all the components shown in the figure.
[0236] From the above description of the embodiments, those skilled in the art can clearly understand that the present disclosure can be realized by means of software and the necessary general hardware platform. Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in various embodiments or some parts of the embodiments.
[0237] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the relevant parts are referred to the part of the method embodiment. The above described system and system embodiment are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to the actual needs, some or all of the modules can be selected to realize the purpose of the embodiment. Those skilled in the art can understand and implement without creative labor.
[0238] The above describes the method for generating a picture and the electronic device provided by the present disclosure in detail. The principles and implementation manners of the present disclosure are described by using specific examples. The above description of the embodiments is only used to help understand the method of the present disclosure and the core idea thereof. Meanwhile, for those skilled in the art, according to the idea of the present disclosure, the specific implementation manners and application ranges can be changed. In conclusion, the present disclosure should not be understood as a limitation of the present disclosure.
Claims
1. A method of generating a picture, wherein, The method comprises: determining user's generation demand information of a picture, the generation demand information at least comprising a subject object image; generating a semantic label for the subject object image by an artificial intelligence (AI) large model, so as to obtain a description text for describing features of the subject object image according to the semantic label, and convert the description text into a first vector; determining elements of multiple types respectively matched with the subject object image according to the first vector corresponding to the subject object image and second vectors respectively corresponding to multiple different types of elements in a pre-established element asset library; wherein the elements are required in a picture generation process, and the second vector is obtained by converting a description text for describing features of the element after generating a semantic label for the element by the AI large model. organizing the subject object image and the matched elements according to a preset picture description protocol to generate a target picture.
2. The method of claim 1, wherein when generating the semantic label for the subject object image, the generated semantic label comprises a category of the subject object, a color tendency, a suitable occasion, a style and / or a crowd.
3. The method of claim 1 or 2, wherein when generating the semantic label for the element, the generated semantic label comprises a suitable scene, a style, a crowd, a category of the subject object, an aspect ratio and / or a quantity of the element.
4. The method of any one of claims 1 to 3, wherein the elements in the element asset library comprise layout scheme elements, content elements, style elements and background elements, wherein the content elements comprise multiple subtypes, the subtypes comprise image types, text types or component types, the style elements are used to describe style configuration information of text and / or components, and the style comprises a font and / or a color matching; wherein the layout scheme elements are used to define a canvas layout template structure, and define relative positions of the content elements in the canvas by a structured relative positioning design method, so as to fill the matched content elements into corresponding positions in the canvas according to definitions of the matched target layout scheme elements.
5. The method of claim 4, wherein the layout scheme elements are used to divide the canvas into multiple information regions, and define subtypes of the content elements required to be filled in each information region and alignment manner information of the content elements in the information region, so as to fill the content elements of the corresponding subtypes into the corresponding information regions according to the alignment manner.
6. The method of claim 5, wherein the alignment manner information of the content elements in the information region comprises: alignment manners of the content elements in a main axis direction and / or an auxiliary axis direction of the information region, the main axis direction being an arrangement direction of the content elements in the information region, and the auxiliary axis direction being a perpendicular direction of the main axis direction.
7. The method of any one of claims 4 to 6, wherein In the picture generation process according to the generation requirement information, for the content of the component class, the size of the background picture frame required to carry the text content and the font size information required to be displayed in the component are determined first, and then the background picture frame of the component is drawn, so that the size of the background picture frame of the component can be adaptively changed with the text displayed in the component.
8. The method of any one of claims 4 to 7, wherein, A plurality of breakpoints are set in advance in the size and / or aspect ratio dimension, and corresponding layout scheme class elements are configured in the element asset library for the plurality of breakpoints; When generating the picture according to the generation requirement information, the target size and / or aspect ratio of the picture to be generated are determined, and if the target size and / or aspect ratio does not hit the breakpoint, the layout scheme class element with the closest size and / or aspect ratio is selected, and the position of the content class element is adaptively determined according to the structured relative positioning design information defined by the layout scheme class element.
9. The method of any one of claims 4 to 8, wherein, In the process of determining the matching element, after the matching background element is determined, the color matching class element is determined by programmatic generation.
10. The method of claim 9, wherein, When the color matching class element is determined by programmatic generation, it includes: Determining the main color of the matched background class element, and defining a set of offsets from the main color as a starting point, and through the offsets, the generated main color system is changed in hue, S saturation and L lightness, and the generated color version color meets the defined requirements.
11. The method of any one of claims 1 to 10, wherein, The generation requirement information further includes text content, which includes text content to be added to the generated picture, and / or text information describing the required scene, style, and required target population.
12. The method of any one of claims 1 to 10, wherein, The subject object image includes a product image, and the generated target picture includes a marketing material picture for participating in a marketing activity.
13. A method of generating a picture, wherein, Including: Receiving user generation requirement information of a picture, the generation requirement information at least including a subject object image; The generated requirement information is submitted to a server, and the server is configured to generate semantic labels for the subject object image by using an AI large model, so as to obtain description text for describing features of the subject object image according to the semantic labels, and then convert the description text into a first vector; according to the first vector corresponding to the subject object image and second vectors corresponding to a plurality of different types of elements in a pre-established element asset library, a plurality of types of elements respectively matched with the subject object image are determined, and the subject object image and the matched elements are organized according to a preset picture description protocol to generate a target picture; wherein the elements are elements required in the generation of the picture, and the second vectors are obtained by converting description text for describing features of the elements after generating semantic labels for the elements by using the AI large model according to the semantic labels; The generated target picture is received and displayed.
14. A computer readable storage medium having stored thereon a computer program, wherein, The program is executed by the processor to implement the steps of the method of any one of claims 1 to 13.
15. An electronic device, comprising: Comprise: One or more processors; And The memory associated with the one or more processors is used to store program instructions, and the program instructions are read and executed by the one or more processors to perform the steps of the method of any one of claims 1 to 13.
16. A computer program product comprising computer program / computer executable instructions which when executed cause a computer to perform the method of claim 15. The computer program / computer executable instructions are executed by the processor in the electronic device to implement the steps of the method of any one of claims 1 to 13. The computer program / computer executable instructions are executed by the processor in the electronic device to implement the steps of the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Advertisement generation method and system, medium and equipment
CN117372087A
Image description text generation method and device based on AIGC, and storage medium
CN117746143A
Image generation method and device, equipment and storage medium
CN118429474A
Method for automatically generating picture and electronic equipment
CN119540376A
Systems and methods for multimodal layout designs of digital publications
US20240104809A1