Poster generation method and device, electronic equipment and storage medium
The poster generation method obtains multimodal poster elements for feature fusion and dynamic layout optimization, solves the problem of low efficiency in the existing technology, and realizes automatic and efficient poster generation.
Patent Information
- Application Number
- CN202510730610.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
Existing poster generation technology has difficulty coordinating content presentation and template adaptation when faced with diverse demands, resulting in low generation efficiency.
By acquiring multimodal poster elements, performing semantic feature extraction and visual feature extraction, combining them with vector conversion for feature fusion, and using the preset poster layout generation model for dynamic layout optimization, the poster is finally rendered to reduce manual intervention.
It realizes the automation of poster generation, improves generation efficiency, balances the conflict between aesthetic rules and content adaptation, and improves the efficiency of poster generation.
Smart Images

Figure CN120672904A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and is applicable to financial technology scenarios and medical technology scenarios, and in particular to a poster generation method and device, electronic equipment and storage medium. Background Art
[0002] Poster generation technology uses algorithms or models to automatically design posters, typically based on template libraries or heuristic rules to achieve personalized visual creation. Poster generation technology can be applied in a variety of scenarios. For example, in fintech, it can be used to generate marketing materials for product promotions and agent publicity. In medical technology, it can be used to generate promotional materials for medical knowledge, public health, and other areas.
[0003] Currently, poster generation technology mainly relies on manually designed heuristic rules or template libraries. However, when faced with diverse poster needs, it is difficult to coordinate factors such as content presentation and template adaptation, which affects the efficiency of poster generation. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a poster generation method and device, an electronic device and a storage medium, aiming to improve the efficiency of poster generation.
[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a poster generation method, the method comprising:
[0006] Acquire a multimodal poster element; wherein the multimodal poster element includes a text element, an image element, and a graphic element;
[0007] Extracting semantic features from the text elements to obtain original text features;
[0008] Extracting visual features of the image elements to obtain original image features;
[0009] Performing vector conversion on the graphic elements to obtain original graphic features;
[0010] Performing feature fusion on the original text features, the original image features, and the original graphic features to obtain original poster fusion features;
[0011] Performing poster layout on the original poster fusion features through a preset poster layout generation model to generate target poster layout features;
[0012] Poster rendering is performed based on the target poster layout features to obtain a target poster.
[0013] In some embodiments, the poster layout generation model includes an original gating network and an original expert network; performing poster layout on the original poster fusion features using a preset poster layout generation model to generate target poster layout features includes:
[0014] Performing gated decision on the original poster fusion features through the original gating network to obtain expert network decision features;
[0015] Based on the decision features of the expert network, the original expert network is constructed to obtain a target expert network;
[0016] The target expert network is used to perform layout prediction on the original poster fusion features to obtain the target poster layout features.
[0017] In some embodiments, the expert network decision feature includes an expert network weight feature, and constructing the original expert network based on the expert network decision feature to obtain a target expert network includes:
[0018] Determining expert network selection information based on the expert network weight characteristics;
[0019] The original expert network is screened based on the expert network selection information to obtain the target expert network.
[0020] In some embodiments, the target expert network includes multiple target expert sub-networks; performing layout prediction on the original poster fusion features through the target expert network to obtain the target poster layout features includes:
[0021] Performing layout reasoning on the original poster fusion features through the target expert sub-network to obtain poster layout sub-features;
[0022] The poster layout sub-features output by each target expert sub-network are weighted and summed based on the expert network weight feature to obtain the target poster layout feature.
[0023] In some embodiments, performing a gating decision on the original poster fusion features by the original gating network to obtain an expert network decision feature includes:
[0024] Performing a linear transformation on the original poster fusion features to obtain an original weight matrix;
[0025] Performing matrix element screening on the original weight matrix to obtain an initial weight matrix;
[0026] The initial weight matrix is normalized to obtain the expert network decision feature.
[0027] In some embodiments, the original weight matrix includes a plurality of original matrix elements; and the step of screening the matrix elements of the original weight matrix to obtain the initial weight matrix includes:
[0028] Sort the original matrix elements based on the values of the original matrix elements to obtain an original element sequence;
[0029] Element screening is performed on the original element sequence to obtain the initial weight matrix.
[0030] In some embodiments, the performing feature fusion on the original text features, the original image features, and the original graphic features to obtain the original poster fusion features includes:
[0031] Performing mapping transformation on the original image features and the original graphic features to obtain a target mapping matrix; wherein the target mapping matrix includes a key mapping matrix and a value mapping matrix;
[0032] Performing aggregation calculation on the original text features and the key mapping matrix to obtain a fused attention score;
[0033] Normalizing the fused attention score to obtain a fused attention weight;
[0034] The fused attention weight and the value mapping matrix are aggregated and calculated to obtain the original poster fusion feature.
[0035] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a poster generation device, comprising:
[0036] A poster element acquisition module, configured to acquire multimodal poster elements, wherein the multimodal poster elements include text elements, image elements, and graphic elements;
[0037] A semantic feature extraction module is used to extract semantic features from the text elements to obtain original text features;
[0038] A visual feature module is used to extract visual features of the image elements to obtain original image features;
[0039] An image feature conversion module, configured to perform vector conversion on the graphic elements to obtain original graphic features;
[0040] A poster feature fusion module is used to fuse the original text features, the original image features and the original graphic features to obtain original poster fusion features;
[0041] A poster layout module is used to perform poster layout on the original poster fusion features through a preset poster layout generation model to generate target poster layout features;
[0042] The poster rendering module is used to perform poster rendering based on the target poster layout features to obtain the target poster.
[0043] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0044] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0045] The poster generation method and device, electronic device, and storage medium proposed in this application obtain multimodal poster elements; wherein the multimodal poster elements include text elements, image elements, and graphic elements, and then perform semantic feature extraction on the text elements to obtain original text features, perform visual feature extraction on the image elements to obtain original image features, and perform vector conversion on the graphic elements to obtain original graphic features, thereby obtaining a unified embedded representation. Next, feature fusion is performed on the original text features, original image features, and original graphic features to obtain original poster fusion features, which solves the problem that traditional template libraries have difficulty coordinating multi-element matching. Furthermore, the original poster fusion features are subjected to poster layout through a preset poster layout generation model to generate target poster layout features, which can dynamically optimize the layout of the poster and effectively balance the conflict between aesthetic rules and content adaptation. Finally, the target poster layout features are subjected to poster rendering to obtain the target poster, realizing automated poster generation, reducing manual intervention links, and improving the efficiency of poster generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flowchart of the poster generation method provided in an embodiment of the present application;
[0047] Figure 2 yes Figure 1 Flowchart of step S105 in FIG.
[0048] Figure 3 yes Figure 1 Flowchart of step S106 in FIG.
[0049] Figure 4 yes Figure 3 Flowchart of step S301 in FIG.
[0050] Figure 5 yes Figure 4 Flowchart of step S402 in FIG.
[0051] Figure 6 yes Figure 3 Flowchart of step S302 in FIG.
[0052] Figure 7 yes Figure 3 Flowchart of step S303 in FIG.
[0053] Figure 8 It is a structural diagram of a poster generating device provided in an embodiment of the present application;
[0054] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0056] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0058] First, let’s analyze some of the terms used in this application:
[0059] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0060] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.
[0061] Poster generation, poster generation technology uses computer algorithms or models to automatically or semi-automatically design posters, usually combining image processing, template libraries, artificial intelligence (AI), user input parameters, and heuristic rules to achieve personalized visual creation. The core of poster generation technology is to quickly generate graphic and text content that meets the requirements of the theme through preset templates, intelligent typesetting, style transfer, or generative adversarial networks (GAN) and other technologies. This technology can be divided into two categories: one is based on a rule system (such as a drag-and-drop design platform), and the other relies on deep learning (such as AI generating original images based on text descriptions). Poster design technology can lower the design threshold, improve efficiency, and support personalized customization. It is widely used in marketing, event promotion, social media and other fields. Some advanced systems can also analyze user preferences or market trends to optimize design effects.
[0062] Gating Network: A gating network is a mechanism for dynamically selecting or weighting multiple submodules (such as expert networks), and is commonly found in models such as mixture of experts (MoE). The core function of a gating network is to generate a set of weights or probability distributions based on the characteristics of the input data to determine the degree of participation of different expert networks. Gating networks are usually implemented by lightweight neural networks (such as fully connected layers or attention mechanisms), which can achieve efficient routing at a low computational cost. For example, in natural language processing, gating networks can learn the preferences of different lexicons or syntactic structures for specific experts, thereby achieving conditional computation. Its advantages lie in flexibility and scalability, and it can increase model capacity without significantly increasing the computational burden.
[0063] Expert Networks (ANs) are specialized sub-models within a hybrid expert system, where each expert typically optimizes for a specific pattern or domain of input data. In the Mixture of Experts (MoE) framework, multiple expert networks exist in parallel, each handling subtasks assigned by a gating network. Experts can be homogeneous (e.g., neural networks with identical structures) or heterogeneous (e.g., models with different architectures), and their design depends on the task requirements. For example, in a vision task, different experts might process texture, color, or shape features, respectively. The outputs of the expert networks are combined through gating weights to form the final prediction. This division of labor can improve the model's specialization while reducing computational cost through sparse activation.
[0064] Poster generation technology can be applied to a variety of application scenarios. For example, in the financial technology scenario, poster generation technology can be used to generate marketing materials such as product promotion and agent promotion; in the medical technology scenario, poster generation technology can be used to generate promotional materials such as medical knowledge and public health.
[0065] Currently, poster generation technology mainly relies on manually designed heuristic rules or template libraries. However, when faced with diverse poster needs, it is difficult to coordinate factors such as content presentation and template adaptation, which affects the efficiency of poster generation.
[0066] In recent years, deep learning technologies, such as LayoutGAN based on generative adversarial networks (GAN) and LayoutVAE based on variational autoencoders (VAE), can also achieve poster generation through end-to-end training. However, deep learning technologies still have defects. On the one hand, the traditional mixture of experts (MoE) model independently routes in the decoding stage and does not fully utilize the semantic priors output by the encoder, resulting in a low match between the activation of the expert network and the layout requirements, affecting the quality of poster generation; on the other hand, the encoder-decoder model with a fixed architecture performs complete calculations on all inputs, resulting in resource waste in simple layout scenarios, affecting the efficiency of poster generation.
[0067] Based on this, the embodiments of the present application provide a poster generation method and device, an electronic device, and a storage medium, aiming to improve the efficiency of poster generation.
[0068] The poster generation method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the poster generation method in the embodiments of the present application is described.
[0069] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0070] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0071] The poster generation method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The poster generation method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the poster generation method, etc., but is not limited to the above forms.
[0072] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0073] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0074] Figure 1 This is an optional flowchart of the poster generation method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107.
[0075] Step S101, obtaining multimodal poster elements; wherein the multimodal poster elements include text elements, image elements and graphic elements;
[0076] Step S102: extracting semantic features from text elements to obtain original text features;
[0077] Step S103, extracting visual features of image elements to obtain original image features;
[0078] Step S104, performing vector conversion on the graphic elements to obtain original graphic features;
[0079] Step S105, performing feature fusion on the original text features, the original image features, and the original graphic features to obtain the original poster fusion features;
[0080] Step S106, performing poster layout on the original poster fusion features using a preset poster layout generation model to generate target poster layout features;
[0081] Step S107 , performing poster rendering based on the target poster layout features to obtain the target poster.
[0082] Steps S101 to S107 shown in the embodiment of the present application are performed by obtaining multimodal poster elements; wherein the multimodal poster elements include text elements, image elements and graphic elements, and then semantic feature extraction is performed on the text elements to obtain original text features, visual feature extraction is performed on the image elements to obtain original image features, and vector conversion is performed on the graphic elements to obtain original graphic features, thereby obtaining a unified embedded representation. Next, feature fusion is performed on the original text features, original image features and original graphic features to obtain original poster fusion features, which solves the problem that traditional template libraries are difficult to coordinate multi-element matching. Furthermore, the original poster fusion features are subjected to poster layout through a preset poster layout generation model to generate target poster layout features, which can dynamically optimize the layout of the poster and effectively balance the conflict between aesthetic rules and content adaptation. Finally, poster rendering is performed on the target poster layout features to obtain the target poster, thereby realizing automated poster generation, reducing manual intervention links, and improving the efficiency of poster generation.
[0083] In step S101 of some embodiments, the multimodal poster elements are manually input by business personnel, designers, and other personnel. The multimodal poster elements are the elements that need to be displayed on the poster, which can ensure that the basic materials for poster generation are complete, making the poster richer in visual and semantics.
[0084] Multimodal poster elements include text elements, image elements, and graphic elements. Specifically, text elements are the textual content in a poster, such as titles, slogans, and descriptions; image elements are the visual elements in a poster, such as photos, illustrations, and background images; and graphic elements are non-photographic design elements in a poster, such as logos, geometric shapes, icons, and decorative lines.
[0085] For example, in a fintech scenario, to generate a promotional poster for a life insurance product, the multimodal poster elements used include:
[0086] Text elements: Title: "Protecting the future, ensuring a lifetime of peace of mind - XX Life Insurance, providing dual protection for you and your family"; Product selling points: "Maximum coverage of xxx, covering major illnesses", "Quick claims settlement"; Call to action: "Consult now to receive your exclusive protection plan";
[0087] Graphic elements: The background image is a soft blue gradient, the main visual image is a warm family photo, and the product diagram is a visual diagram of the insurance coverage;
[0088] Graphic elements: insurance company's logo, background frame.
[0089] In the medical technology scenario, to generate a promotional poster about vaccination, the multimodal poster elements used include:
[0090] Text elements: Main title: "Protecting health starts with scientific prevention", subtitle: "Influenza vaccination initiative"; body content: "Influenza vaccination can reduce the risk of severe illness and protect you and your family", "Recommended population for vaccination: the elderly, children, patients with chronic diseases, etc.", "Vaccination location: xx Community Hospital", "Consultation phone number: xxxx";
[0091] Graphic elements: The background image is a soft blue gradient background, the main visual image is a photo showing medical staff vaccinating an elderly person, and the illustration is a cartoon-style diagram of the virus and antibodies;
[0092] Graphic elements: community hospital or public health department logo, background frame.
[0093] In some embodiments, the image elements and graphic elements may be manually screened from a template library, or may be manually acquired from other channels, without limitation thereto.
[0094] In step S102 of some embodiments, natural language processing techniques (e.g., BERT model, Word2Vec model, etc.) are used to extract semantic features from text elements to obtain structured data, i.e., raw text features, which help improve the semantic coherence and personalized expression of the poster. Specifically, the raw text features are word vectors corresponding to the text.
[0095] In step S103 of some embodiments, a pre-trained visual model (such as a CNN model, a ViT model, etc.) is used to analyze and extract visual features such as color, texture, object, style, etc. of image elements, convert the image elements into high-dimensional features, and obtain the original image features to facilitate subsequent poster layout.
[0096] In step S104 of some embodiments, the graphic elements are converted into a vector format, namely, the SVG format, to form structured original graphic features, which facilitates subsequent poster layout.
[0097] It should be noted that the original text features, original image features, and original graphic features have the same data dimension, thereby achieving cross-modal alignment and ensuring semantic consistency, which can be used for subsequent multimodal fusion.
[0098] See also Figure 2 In some embodiments, step S105 may include but is not limited to steps S201 to S204:
[0099] Step S201, mapping and transforming the original image features and the original graphic features to obtain a target mapping matrix; wherein the target mapping matrix includes a key mapping matrix and a value mapping matrix;
[0100] Step S202: performing aggregation calculation on the original text features and the key mapping matrix to obtain a fused attention score;
[0101] Step S203, normalize the fused attention score to obtain the fused attention weight;
[0102] Step S204: perform aggregation calculation on the fused attention weight and value mapping matrix to obtain the original poster fusion feature.
[0103] In the steps S201 to S204 shown in the embodiment of the present application, a key mapping matrix and a value mapping matrix are obtained by mapping and transforming the original image features and the original graphic features, thereby establishing a cross-modal interaction foundation. Next, the original text features and the key mapping matrix are aggregated and calculated to obtain a fused attention score, thereby realizing content-oriented visual element screening and ensuring that key information is accurately focused. Furthermore, the fused attention score is normalized and calculated to obtain a fused attention weight, which can reflect which image information is more critical in the poster generation process. Finally, the fused attention weight and the value mapping matrix are aggregated and calculated to obtain the original poster fusion feature, establish semantic associations between elements of different modalities, capture semantic associations and spatial constraints between elements, and realize deep fusion between features of different modalities, which is conducive to improving the accuracy and efficiency of poster layout.
[0104] In step S201 of some embodiments, a fully connected layer is used to perform a linear transformation on the original image features and the original graphic features, and the original image features and the original graphic features are mapped to the same space to obtain a key mapping matrix and a value mapping matrix, wherein the key mapping matrix is used for similarity calculation and the value mapping matrix is used for feature reconstruction.
[0105] In step S202 of some embodiments, a dot product is calculated between the original text feature and the key mapping matrix to obtain a fused attention score; specifically, the fused attention score is used to measure the degree of association between the elements in the original text feature and the elements in the key mapping matrix.
[0106] In step S203 of some embodiments, the fused attention score can be normalized and calculated using methods such as softmax function, Z-score normalization, and maximum normalization to obtain a fused attention weight, where the fused attention weight is used to represent the degree of association between the elements in the original text features and the elements in the key mapping matrix.
[0107] In step S204 of some embodiments, the original poster fusion features are obtained by performing weighted sum calculation on the fusion attention weights and the value mapping matrix, and semantic associations between different modal elements are established to achieve deep fusion between different modal features.
[0108] It is understandable that the cross-attention mechanism achieves deep fusion of features from different modalities. Compared with traditional poster generation methods, it can adaptively handle diverse input combinations, improving the accuracy and efficiency of content generation.
[0109] In step S106 of some embodiments, the preset poster layout generation model is pre-trained. Specifically, the poster layout generation model can adopt models such as a hierarchical expert model and a sparse hybrid expert model, and can be dynamically routed to a specific expert sub-network through a gating mechanism to achieve refined poster layout.
[0110] In one embodiment, the poster layout generation model adopts a sparse mixture of experts model, and the gating network adopts Top-k selection, retaining only the weights of the first k expert networks, thereby activating the corresponding expert network according to the weight.
[0111] See also Figure 3 In some embodiments, the poster layout generation model includes an original gating network and an original expert network, and the original expert network includes multiple original expert sub-networks; step S106 may include but is not limited to steps S301 to S303:
[0112] Step S301, performing gated decision on the original poster fusion features through the original gated network to obtain expert network decision features;
[0113] Step S302: constructing the original expert network based on the expert network decision features to obtain a target expert network;
[0114] Step S303: Layout prediction is performed on the original poster fusion features through the target expert network to obtain target poster layout features.
[0115] In steps S301 to S303, as shown in the embodiment of this application, the original gated network performs a gated decision on the original poster fusion features to obtain the expert network decision features. Based on the expert network decision features, the original expert network is screened and network constructed to obtain the target expert network. This achieves adaptive optimization of the network structure and can dynamically adjust the expert network processing strategy to meet different poster design requirements. Finally, adaptive optimization of the network structure is achieved, enabling the model to dynamically adjust the processing strategy to meet different design requirements, resulting in a poster layout that meets the poster design requirements, effectively improving the efficiency of poster layout, and thus improving the efficiency of poster generation.
[0116] It should be noted that the model structure of each original expert sub-network can be the same or different, and the specific setting needs to be combined with the actual application scenario, but is not limited to this.
[0117] See also Figure 4In some embodiments, step S301 may include but is not limited to steps S401 to S403:
[0118] Step S401, performing a linear transformation on the original poster fusion features to obtain an original weight matrix;
[0119] Step S402, screening the matrix elements of the original weight matrix to obtain an initial weight matrix;
[0120] Step S403: normalize the initial weight matrix to obtain expert network decision features.
[0121] In steps S401 to S403 shown in the embodiment of the present application, the original weight matrix of each original expert sub-network is obtained by linearly transforming the original poster fusion features. Then, the matrix elements of the original weight matrix are screened, and only the weights of part of the original expert sub-network are retained to obtain the initial weight matrix. Finally, the initial weight matrix of the weights of the remaining original expert sub-network is normalized to achieve sparse processing and obtain the expert network decision feature. The expert network decision feature not only retains the discriminability of the original poster fusion feature, but also highlights the task-related features through the gating mechanism, thereby improving the feature selection ability of the model, and enabling the expert network to adaptively focus on the most effective feature combination, thereby improving the efficiency and accuracy of the poster layout.
[0122] In step S401 of some embodiments, a linear transformation layer or an MLP layer may be used to perform a linear transformation on the original poster fusion features to generate an original weight matrix for each original expert sub-network. The original weight matrix includes multiple original matrix elements, each of which represents the weight of an original expert sub-network.
[0123] See also Figure 5 In some embodiments, step S402 may also include but is not limited to steps S501 to S502:
[0124] Step S501, sorting the original matrix elements based on the values of the original matrix elements to obtain an original element sequence;
[0125] Step S502: Filter the original element sequence to obtain an initial weight matrix.
[0126] In steps S501 to S502 shown in the embodiment of the present application, the original matrix elements are sorted based on the numerical values of the original matrix elements to obtain the original element sequence, and the elements of the original element sequence are screened to obtain the initial weight matrix. The most relevant expert subnetwork can be dynamically selected to achieve adaptive allocation of computing resources, so that the model can better adapt to different poster design requirements, thereby improving the efficiency and accuracy of poster layout.
[0127] It should be noted that the sorting can be performed from largest to smallest by the numerical value of the original matrix elements, or from smallest to largest, without limitation, as long as the sorting can be achieved. This sorting process helps to more quickly find the top-K original matrix elements. K is a preset expert number threshold, which depends on the actual scenario setting, the actual poster design requirements, and the multimodal poster elements, without limitation.
[0128] In step S502 of some embodiments, the top-K elements are screened out from the original element sequence based on a preset expert quantity threshold to obtain an initial weight matrix, wherein the values of the selected original matrix elements are retained, while the values of the unselected original matrix elements are set to negative infinity or zero.
[0129] It should be noted that the initial weight matrix includes multiple initial matrix elements, each of which represents the weight of an original expert sub-network.
[0130] In step S403 of some embodiments, the initial weight matrix is normalized by a Softmax function to obtain the expert network decision feature, thereby ensuring that the weighted combination of the expert sub-network outputs is reasonable and stable.
[0131] For example: assuming that the original expert network includes 6 original expert sub-networks, the original weight matrix includes 6 original matrix elements;
[0132] Assume the original weight matrix is [1.3, 0.8, 0.9, 1.5, 0.8, 0.5];
[0133] Specifically, the value of the original matrix element corresponding to the first original expert sub-network is 1.3,
[0134] The value of the original matrix element corresponding to the second original expert sub-network is 0.8.
[0135] The value of the original matrix element corresponding to the third original expert sub-network is 0.9.
[0136] The value of the original matrix element corresponding to the fourth original expert sub-network is 1.5.
[0137] The value of the original matrix element corresponding to the fifth original expert sub-network is 0.8.
[0138] The value of the original matrix element corresponding to the sixth original expert sub-network is 0.5.
[0139] Sorting the original matrix elements from large to small based on their numerical values, the resulting original element sequence is {1.5, 1.3, 0.9, 0.8, 0.8, 0.5};
[0140] Assume that K in Top-K is set to 2, that is, two elements need to be filtered out from the original element sequence.
[0141] Therefore, a Top-2 screening is performed from the original element sequence {1.5, 1.3, 0.9, 0.8, 0.8, 0.5}, and the two largest values are 1.5 and 1.3. The corresponding expert sub-networks are: the fourth original expert sub-network and the first original expert sub-network. Therefore, the fourth original expert sub-network and the first original expert sub-network retain the values of the original matrix elements.
[0142] The values of the corresponding matrix elements of the remaining third original expert sub-network, second original expert sub-network, fifth original expert sub-network, and sixth original expert sub-network are set to 0, and the initial weight matrix is [1.3, 0, 0, 1.5, 0, 0].
[0143] Furthermore, after normalizing the initial weight matrix [1.3, 0, 0, 1.5, 0, 0] using the Softmax function, the decision feature of the expert network is obtained as [0.452, 0, 0, 0.548, 0, 0].
[0144] It should be noted that the expert network decision features include expert network weight features, and each expert network weight feature corresponds to the weight of an original expert sub-network.
[0145] For example: in the expert network decision feature [0.452, 0, 0, 0.548, 0, 0], the elements 0.452, 0, 0, 0.548, 0, 0 are the expert network weight features.
[0146] The expert network decision feature [0.452, 0, 0, 0.548, 0, 0] contains the following information:
[0147] The number of original expert sub-networks that need to be activated is 2, namely the first original expert sub-network and the fourth original expert sub-network. The weight value of the first original expert sub-network is 0.452, and the weight value of the second original expert sub-network is 0.548.
[0148] See also Figure 6In some embodiments, step S302 includes but is not limited to steps S601 to S602:
[0149] Step S601, determining expert network selection information based on expert network weight characteristics;
[0150] Step S602 : performing network screening on the original expert network based on the expert network selection information to obtain a target expert network.
[0151] In steps S601 to S602 shown in the embodiment of the present application, expert network selection information is determined based on the expert network weight characteristics, and the original expert network is screened and activated based on the expert network selection information to obtain a target expert network, thereby selecting different expert networks according to the multimodal poster characteristics. The selected expert networks work together to achieve adaptive allocation of computing resources, which helps to improve the efficiency and accuracy of poster layout.
[0152] In step S601 of some embodiments, expert network selection information is determined based on the value of each expert network weight feature in the expert network decision feature. Specifically, if the value of the expert network weight feature is not zero, the original expert sub-network is selected; if the value of the expert network weight feature is zero, the original expert sub-network is not selected.
[0153] Next, based on the expert network selection information, the multiple original expert sub-networks in the original expert network are screened to determine the selected original expert sub-network to obtain the target expert network. The target expert network includes multiple target expert sub-networks, each of which is a selected original expert sub-network.
[0154] See also Figure 7 In some embodiments, step S303 may include but is not limited to steps S701 to S702:
[0155] Step S701, performing layout reasoning on the original poster fusion features through the target expert sub-network to obtain poster layout sub-features;
[0156] Step S702 : performing weighted summation on the poster layout sub-features output by each target expert sub-network based on the expert network weight feature to obtain the target poster layout feature.
[0157] In steps S701 to S702 shown in the embodiment of the present application, different target expert sub-networks are used to perform layout inference on the original poster fusion features, infer the local or structural features of the poster layout, improve the feature diversity, and obtain multiple poster layout sub-features; then, based on the expert network weight features, the poster layout sub-features output by each target expert sub-network are weighted and summed to obtain the target poster layout features, thereby improving the quality and efficiency of the poster layout and being more in line with the actual poster generation needs.
[0158] After determining the target expert sub-network, the original poster fusion features are input into the target expert sub-network respectively, so that each target expert sub-network can perform poster layout inference according to its own capabilities, and thus output the corresponding poster layout sub-features.
[0159] Furthermore, based on the expert network weight features output by the original gating network, the poster layout sub-features output by each target expert sub-network are weighted and summed to achieve dynamic fusion of the outputs of different expert sub-networks and adaptively combine the most relevant features, thereby improving the quality and flexibility of poster layout.
[0160] In step S107 of some embodiments, poster rendering is performed on the target poster layout features using a differentiable rendering model to obtain the target poster. The differentiable rendering model may be a PyTorch3D model, a SoftRas model, etc., but is not limited thereto.
[0161] In some embodiments, after step S107, the differentiable rendering model also evaluates the target poster to obtain an aesthetic evaluation result (such as alignment energy and balance score). Furthermore, the aesthetic evaluation result is fed back to the poster layout generation model, so that the poster layout generation model is optimized based on the aesthetic evaluation result, achieving end-to-end joint optimization.
[0162] The poster generation method provided in the embodiments of this application uses a layered visual language Transformer encoder to process multimodal input elements (text, images, and graphics), generates a context-aware feature matrix, and introduces a sparse mixture of experts (SMoE) decoder for dynamic routing. Compared with traditional rule-based poster generation methods, this method can respond to diverse needs more quickly and reduce computational overhead by 40%-60%. It is particularly suitable for commercial application environments that require high-frequency, multi-scene poster output, significantly improving the efficiency and quality of poster layout generation.
[0163] In addition, the poster generation method provided in the embodiments of the present application can be applied to multimodal content generation scenarios such as illustration design, advertising design, web design, and UI / UX design. Furthermore, in addition to being applied to financial technology and medical technology scenarios, the poster generation method can also be used in other application scenarios, such as retail and e-commerce scenarios, education scenarios, and public service promotion scenarios, without limitation.
[0164] See also Figure 8 The embodiment of the present application further provides a poster generation device that can implement the above poster generation method, and the device includes:
[0165] The poster element acquisition module 801 is used to acquire multimodal poster elements; wherein the multimodal poster elements include text elements, image elements and graphic elements;
[0166] Semantic feature extraction module 802, used to extract semantic features from text elements to obtain original text features;
[0167] The visual feature module 803 is used to extract visual features of image elements to obtain original image features;
[0168] The image feature conversion module 804 is used to perform vector conversion on the graphic elements to obtain the original graphic features;
[0169] The poster feature fusion module 805 is used to fuse the original text features, the original image features and the original graphic features to obtain the original poster fusion features;
[0170] The poster layout module 806 is used to perform poster layout on the original poster fusion features through a preset poster layout generation model to generate target poster layout features;
[0171] The poster rendering module 807 is used to perform poster rendering based on the target poster layout features to obtain the target poster.
[0172] The specific implementation of the poster generating device is substantially the same as the specific embodiment of the poster generating method described above, and will not be described in detail here.
[0173] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above poster generation method. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0174] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0175] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0176] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the poster generation method of the embodiments of this application;
[0177] Input / output interface 903, used to implement information input and output;
[0178] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0179] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0180] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0181] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned poster generation method is implemented.
[0182] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0183] The poster generation method and device, electronic device and storage medium provided in the embodiments of the present application obtain multimodal poster elements; wherein the multimodal poster elements include text elements, image elements and graphic elements, and then perform semantic feature extraction on the text elements to obtain original text features, perform visual feature extraction on the image elements to obtain original image features, and perform vector conversion on the graphic elements to obtain original graphic features, thereby obtaining a unified embedded representation. Then, feature fusion is performed on the original text features, original image features and original graphic features to obtain original poster fusion features, which solves the problem that traditional template libraries are difficult to coordinate multi-element matching. Furthermore, the original poster fusion features are subjected to poster layout through a preset poster layout generation model to generate target poster layout features, which can dynamically optimize the layout of the poster and effectively balance the conflict between aesthetic rules and content adaptation. Finally, the target poster layout features are subjected to poster rendering to obtain the target poster, thereby realizing automated poster generation, reducing manual intervention links, and improving the efficiency of poster generation.
[0184] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0185] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0187] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0188] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0189] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0190] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0191] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0194] The non-Company software tools or components appearing in the embodiments of this application are for illustrative purposes only and do not represent actual use.
[0195] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A poster generation method, characterized in that: The method comprises: Acquire a multimodal poster element; wherein the multimodal poster element includes a text element, an image element, and a graphic element; Extracting semantic features from the text elements to obtain original text features; Extracting visual features of the image elements to obtain original image features; Performing vector conversion on the graphic elements to obtain original graphic features; Performing feature fusion on the original text features, the original image features, and the original graphic features to obtain original poster fusion features; Performing poster layout on the original poster fusion features through a preset poster layout generation model to generate target poster layout features; Poster rendering is performed based on the target poster layout features to obtain a target poster.
2. The method according to claim 1, characterized in that The poster layout generation model includes an original gated network and an original expert network; performing poster layout on the original poster fusion features by using a preset poster layout generation model to generate target poster layout features includes: Performing gated decision on the original poster fusion features through the original gating network to obtain expert network decision features; Based on the decision features of the expert network, the original expert network is constructed to obtain a target expert network; The target expert network is used to perform layout prediction on the original poster fusion features to obtain the target poster layout features.
3. The method according to claim 2, characterized in that The expert network decision feature includes an expert network weight feature, and the network construction of the original expert network based on the expert network decision feature to obtain a target expert network includes: Determining expert network selection information based on the expert network weight characteristics; The original expert network is screened based on the expert network selection information to obtain the target expert network.
4. The method according to claim 3, characterized in that The target expert network includes a plurality of target expert sub-networks; performing layout prediction on the original poster fusion features through the target expert network to obtain the target poster layout features includes: Performing layout reasoning on the original poster fusion features through the target expert sub-network to obtain poster layout sub-features; The poster layout sub-features output by each target expert sub-network are weighted and summed based on the expert network weight feature to obtain the target poster layout feature.
5. The method according to claim 2, characterized in that The gated decision is performed on the original poster fusion features by the original gating network to obtain the expert network decision features, including: Performing a linear transformation on the original poster fusion features to obtain an original weight matrix; Performing matrix element screening on the original weight matrix to obtain an initial weight matrix; The initial weight matrix is normalized to obtain the expert network decision feature.
6. The method according to claim 5, characterized in that The original weight matrix includes a plurality of original matrix elements; The performing matrix element screening on the original weight matrix to obtain an initial weight matrix includes: Sort the original matrix elements based on the values of the original matrix elements to obtain an original element sequence; Element screening is performed on the original element sequence to obtain the initial weight matrix.
7. The method according to any one of claims 1 to 6, characterized in that The step of fusing the original text features, the original image features, and the original graphic features to obtain the original poster fusion features includes: Performing mapping transformation on the original image features and the original graphic features to obtain a target mapping matrix; wherein the target mapping matrix includes a key mapping matrix and a value mapping matrix; Performing aggregation calculation on the original text features and the key mapping matrix to obtain a fused attention score; Normalizing the fused attention score to obtain a fused attention weight; The fused attention weight and the value mapping matrix are aggregated and calculated to obtain the original poster fusion feature.
8. A poster generating device, characterized in that: The device comprises: A poster element acquisition module, configured to acquire multimodal poster elements, wherein the multimodal poster elements include text elements, image elements, and graphic elements; A semantic feature extraction module is used to extract semantic features from the text elements to obtain original text features; A visual feature module is used to extract visual features of the image elements to obtain original image features; An image feature conversion module, configured to perform vector conversion on the graphic elements to obtain original graphic features; A poster feature fusion module is used to fuse the original text features, the original image features and the original graphic features to obtain original poster fusion features; A poster layout module is used to perform poster layout on the original poster fusion features through a preset poster layout generation model to generate target poster layout features; The poster rendering module is used to perform poster rendering based on the target poster layout features to obtain the target poster.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.