Reinforcement learning-based banner synthesis method and apparatus
Through the method based on reinforcement learning, the problem of high manual intervention in UI material size transformation is solved, automated processing and layer combination optimization are realized, and design efficiency and resource utilization are improved.
Patent Information
- Application Number
- PCT/CN2024/138796
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art has a high degree of manual intervention in the size transformation of UI materials, which leads to inefficiency in adapting booths, versions, and models of different sizes, and is prone to waste manpower, making it difficult to optimize layer combinations.
The banner synthesis method based on reinforcement learning is adopted, and the optimized model is generated through layer preprocessing, multimodal feature extraction, dynamic sampling and reinforcement learning training to realize automated size transformation and layer combination optimization.
It realizes automation of UI material size transformation, reduces manual intervention, improves design efficiency, is versatile and scalable, and can better utilize designer time and resources.
Smart Images

Figure CN2024138796_19062025_PF_FP_ABST
Abstract
Description
A banner synthesis method and device based on reinforcement learning
[0001] Related applications
[0002] This application claims priority to Chinese patent application No. 2023117153925, filed on December 14, 2023, entitled “A banner synthesis method and device based on reinforcement learning,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application belongs to the field of IT technology, and specifically relates to a banner synthesis method and device based on reinforcement learning. Background Art
[0004] In the field of visual design, designers often spend considerable time on simple tasks, such as revising copy, designing simple poster layouts, and expanding them to multiple sizes for different models and booths. This work consumes a significant amount of time and labor (5-6 posters per person-day), but its contribution to the designer's advancement and growth is very limited. On the other hand, precision marketing is a major trend in the future. In the context of high traffic, homepage poster resource booths need to present a personalized effect, which also places very high demands on poster production efficiency.
[0005] After designing a banner, designers often spend a significant amount of time adapting it to different sizes, for different booths, versions, and models. However, the varying sizes of asset banners for different booths necessitate significant changes to the layout of the embedded layers: scaling, cropping, and expansion. This creates significant challenges for designers. This means that each booth version requires a completely different design, which wastes significant effort. Furthermore, adapting to multiple sizes when the assets are already fixed and their relative positions are nearly certain is essentially an optimization problem. Summary of the Invention
[0006] The purpose of this application is to provide a banner synthesis method based on reinforcement learning, which can be applied to the automatic resizing of materials in APP resource positions and advertising positions, freeing UI designers from monotonous and repetitive tasks such as resizing, liberating the productivity of UI workers, and having universality for layer combination optimization problems, and the model design is scalable.
[0007] The technical solutions adopted in this application are as follows:
[0008] A banner synthesis method based on reinforcement learning, comprising:
[0009] S1. Preprocess the sample layer and perform sample expansion, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping;
[0010] S2. Extract multimodal features from the expanded samples through the encoder, represent the multimodal features with feature vectors, and enhance the generalization ability of the model through multimodal feature learning. The multimodal features include structural features, design text features, and visual features.
[0011] S3. Solve the biased training problem for the multimodal samples and use dynamic sampling to generate reinforcement learning training samples online;
[0012] S4. Establish a model training layer, train the training model, and generate an optimized model;
[0013] S5. Provide services to users through the optimized model in the form of a Restful interface.
[0014] In a preferred embodiment, the pre-processing step of layer alignment, merging and cropping from the sample includes:
[0015] Get an approximate KV and filter the artboards to determine which artboard is the primary KV and align the layers.
[0016] Scale each layer based on the size ratio between the main KV and the artboard to be processed.
[0017] In a preferred embodiment, the step of expanding the sample includes:
[0018] Obtain available main KVs from the sample so that each main KV has several label data of different sizes on average;
[0019] The above main KVs are paired in a pairwise manner to generate samples, that is, to obtain enhanced sample data.
[0020] In a preferred embodiment, the encoder adopts a hybrid DNN network design feature to encode the extracted multimodal features in a DNN manner, wherein the multimodal features are represented by feature vectors and the multimodal features need to be mapped to a unified vector space.
[0021] In a preferred solution, the biased training is performed by designing a replayBuffer mechanism, and the dynamic sampling method adopts a random sampling method that combines uniform sampling and optimized sampling.
[0022] In a preferred solution, the model training layer maps the elements of the reinforcement learning framework such as the agent's state and observation information to the vector space, uses the Qmix network for learning, and designs a padding algorithm to solve the problem of dynamic agent training.
[0023] In a preferred embodiment, the Qmix network is used for learning using a centralized training distributed usage architecture, specifically comprising the following steps:
[0024] Define a global optimization goal, Qtot, where the Q value needs to satisfy:
[0025] That is, the Q function Qa of each agent is guaranteed to be monotonically increasing;
[0026] Optimize the Q value of each agent. The optimization relationship is as follows:
[0027] Design a mixer subnetwork to obtain the overall Q value and ensure the monotonically increasing function;
[0028] Design the Agent sub-network to optimize and maximize the Q value;
[0029] A GRU network module is also added to the Agent sub-network to capture historical action information.
[0030] In a preferred embodiment, the step of designing a mixer subnetwork includes:
[0031] Design a fusion network, input a series of optimized Q values, and obtain an overall Q value through a function transformation;
[0032] Design a two-layer MLP network and use gradient to ensure the monotonically increasing function.
[0033] In a preferred embodiment, the method for designing a padding algorithm to solve the problem of dynamic agent training includes:
[0034] First, set a maximum number of supported agents N. This number is a hyperparameter, so you can try to find an optimal value in experiments. Each time you input a sample with n layers, only the first n of the N agents are activated, and the rest are inactive.
[0035] The QLearner module is used to train and optimize the agent subnetwork to maximize the Q value of the optimal strategy action.
[0036] Input action through the Environment module to obtain the state, observation and corresponding reward caused by the current action.
[0037] A banner synthesis device based on reinforcement learning is applied to the above-mentioned banner synthesis method based on reinforcement learning. The banner synthesis device based on reinforcement learning includes:
[0038] A preprocessing module is used to perform layer preprocessing on the sample and expand the sample, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping;
[0039] The feature module is used to extract multimodal features from the expanded samples and represent the multimodal features with feature vectors. The multimodal features include structural features, design text features, and visual features.
[0040] The sample enhancement module is used to solve the biased training problem of multimodal samples and generate reinforcement learning training samples online using dynamic sampling.
[0041] Model training module, used to train the model and generate an optimized model;
[0042] The service module provides services to users through the form of a Restful interface after optimizing the model. The effect achieved by this application is that the current industry's scale transformation of UI materials is mostly based on manual rules and simple naive classifier algorithms, with a very high degree of human intervention. This application uniquely introduces distributed multi-agent reinforcement learning, combined with dynamic random sampling, and uses the reinforcement learning algorithm framework to model the scale change optimization problem, solving the scale transformation problem in a fully automated end-to-end manner.
[0043] This application innovatively combines an agent feature extraction network structure to fully learn the agent's multimodal features and optimize model learning. In addition, during the sample construction phase, a layer alignment algorithm and a noise reduction algorithm are proposed to give the model stronger generalization capabilities and enrich the learning samples.
[0044] While industry practices for multi-agent reinforcement learning typically use a fixed number of agents, this scenario requires a large and dynamic number of agents. We designed a gradient masking algorithm to address dynamic padding in tensor calculations and a merge mapping algorithm to reduce computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0046] FIG1 is a flow chart of the present application;
[0047] Figure 2 is a framework diagram of this application;
[0048] FIG3 is a structural diagram of a feature module in this application;
[0049] FIG4 is a schematic diagram of the structure of the qmixer module in this application;
[0050] FIG5 is a schematic diagram of gradient tensor masking in this application;
[0051] Figure 6 is a schematic diagram of DQN training prediction in this application;
[0052] FIG7 is a modeling architecture of the service module in this application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0055] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present application. The phrase "in a preferred embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive with other embodiments.
[0056] Please refer to FIG. 1 , which shows the first embodiment of the present application. This embodiment provides a banner synthesis method based on reinforcement learning, including:
[0057] S1. Preprocess the sample layer and perform sample expansion, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping;
[0058] S2. Extract multimodal features from the expanded samples through the encoder, represent the multimodal features with feature vectors, and enhance the generalization ability of the model through multimodal feature learning. The multimodal features include structural features, design text features, and visual features.
[0059] S3. Solve the biased training problem for the multimodal samples and use dynamic sampling to generate reinforcement learning training samples online;
[0060] S4. Establish a model training layer, train the training model, and generate an optimized model;
[0061] S5. Provide services to users through the optimized model in the form of a Restful interface.
[0062] The above approach is used because the industry's current resizing of UI assets is mostly based on manual rules and simple naive classifier algorithms, resulting in a very high degree of manual intervention. Therefore, this application innovatively introduces distributed multi-agent reinforcement learning, combined with dynamic random sampling, to model the resizing optimization problem using a reinforcement learning algorithm framework, achieving an end-to-end fully automated solution to the resizing problem. A feature extraction network structure is employed to fully learn the multimodal features of samples and optimize model learning. Furthermore, during the sample construction phase, preprocessing operations such as layer alignment and resizing are proposed to enhance the model's generalization capabilities and enrich the learning samples.
[0063] In a preferred embodiment, the pre-processing step of layer alignment, merging and cropping from the sample includes:
[0064] Get an approximate KV (main visual poster), filter the artboards, determine which artboard is the main KV, and align the layers. This way, we can calculate the reference offset for each layer of the artboard we are processing.
[0065] Scale each layer based on the size ratio between the main KV and the artboard to be processed.
[0066] The step of performing sample expansion includes:
[0067] Obtain available main KVs from the sample so that each main KV has several label data of different sizes on average;
[0068] The above main KVs are paired in a pairwise manner to generate samples, that is, to obtain enhanced sample data.
[0069] In the above method, since the existing samples are handmade works accumulated over time during design, the styles of different designers vary greatly. Therefore, it is necessary to first perform a canvas screening. The screening here can be achieved through a classifier based on features such as size, shape, and aspect ratio. At the same time, a new master KV that is most similar to the original master KV will be selected. The number of layers of this master KV is within a controllable range.
[0070] Layer alignment is primarily based on the principle that, given a single resizing transformation, the layer combinations between artboards of varying sizes will not differ significantly. Therefore, image similarity and other methods can be used for matching and alignment. Since the number of layers is reduced, computational complexity of O(N**2) becomes feasible. After layer alignment, a mapping of all layers to the master KV is obtained. Correspondingly, their size ratios and aspect ratios can be easily calculated. The reference offset for a layer is simply the normalized offset of the layer in the master KV, normalized to between 0 and 1.
[0071] After obtaining each layer of the original KV, the next step is to trim each layer. We found that there is a lot of redundant layer information in the original KV. However, these layers actually only contribute a small part of the information to the finished banner. A large amount of image information is not highlighted due to the shielding of the visual window. This redundant information is of no benefit to our model fitting. Here we will use IOU to first perform the intersection and union calculation of the background window and layer information to extract the most useful key information for each banner. After that, we will perform the scaling calculation of each layer. size =w*h target szie =w′*h′
[0072] As shown in the formula above, we first need to obtain the width (w) and height (h) of the original KV, and then determine the required width (w') and height (h') of the target banner. To ensure that the layer does not deform after the change, we use proportional scaling during the layer size change process.
[0073] There are three ways to determine the scaling ratio: the minimum, maximum, or average of the width-to-height ratios. Here, we choose the maximum, which covers more common scenarios and achieves better results.
[0074] After determining the size transformation ratio of each layer, we use the principle of geometric transformation, supplemented by some rules constructed through manual experience, to calculate the transformation of each layer in the artboard to be processed. Finally, we obtain the transformed layer and use the reference position as the initial position of the layer movement.
[0075] After data cleaning, we selected 500 available main KVs from the original 1000 PSD files. On average, each main KV has 4 different size label data (the result of manual operation). We generate samples by pairing them two by two, which can be obtained.
[0076] With approximately 3,000 available samples, we can expand the original data by more than 3 times after data augmentation to enhance our model training.
[0077] In one embodiment, please refer to Figure 3, the encoder adopts a hybrid DNN network design feature to encode the extracted multimodal features in a DNN manner, wherein, representing the multimodal features with feature vectors requires mapping the multimodal features to a unified vector space.
[0078] In this embodiment, in order to maximize the representation ability of layer features, an encoder parallel structure is adopted, and the original features are all represented as an embedding structure. In order to better model the representation features and enhance the generalization ability of the module, except for the bottom embedding layer, the other parameters of the encoder are all in the form of shared parameters. For the selection of encoder, the features are all encoded in the form of DNN (DNN deep neural network). Firstly, considering that the representation of the features themselves is sufficient, the model does not need to consume too much capacity in the feature layer. Secondly, DNN is relatively simple in engineering implementation and has high enough computational efficiency. For the lowest module of a model, performance efficiency requirements are also relatively important. All features are encoded by their respective encoders and input into the HAL layer (hardware abstraction layer) for state and observation. In addition, the explanation of the above-mentioned special terms will not be emphasized later. Please refer to the explanation here for details.
[0079] Preferably, the present application also discloses that a replayBuffer mechanism is designed for solving biased training, and a dynamic sampling method adopts a random sampling method that combines uniform sampling and optimized sampling.
[0080] Specifically, the feature vector representation of each layer (agent) modeled above cannot be used directly for model training because we must define several elements in the MARL problem. The HAL layer is mainly used for assembly with state and observation. In essence, it assembles the underlying encoded features on demand according to the respective focus of state and observation, and assembles them into new state and observation tensors. In addition, a LayerNorm (normalization) operation is required to normalize the feature distribution differences caused by different samples to ensure stable training.
[0081] The features of an agent (layer) can be abstracted into two categories. One is representational features, which are features that represent the inherent characteristics of the agent, such as color, shape, and reference offset, which are features that do not change throughout the process. The other is not related to the properties of the agent itself, but belongs to a state in the current environment, such as the current coordinates and the current distance from the target position. This vector will change after each step. This distinction makes it clear which ones should be placed in observation and state. Since observation participates in the calculation of the agent's Q value (the Q value recorded in the Q-Table learned by the agent), its value will be iteratively calculated in the GRU module of the agent network. In fact, the calculations of the agents in this network are all parallel and independent, so it is preferred to place the changing features of a single agent, that is, real-time location information. However, state is more biased towards global information because it is used as input in the mixer network. Therefore, it is necessary to reflect the differences between agents in the entire environment. At this time, the state will assemble more representative features of the agent.
[0082] Furthermore, to manually coordinate the common optimization goals of each agent, we also specifically designed the absolute difference between the sum of the inter-agent distances in the sample size palette and the actual sum of the distances (the result of manual operation). Experiments have shown that this significantly improves the stability of agent motion and prevents individual agents from falling into local optimal strategies and drifting far from the target position.
[0083] Among them, all model training requires samples, and MARL (multi-agent reinforcement learning) problems are no exception. However, the samples in RL problems are different from the static samples generated in advance in general supervised learning. All samples must be generated dynamically online during interaction with the environment. Therefore, it is necessary to design a simulator to conduct the interaction between the agent and the environment and collect the episode sample sequence formed by the interaction.
[0084] The most critical entity in the Simulator layer is the actor, which in our scenario is the agent, or layer. During the initial training phase of the model, the parameters of each agent network are still randomly initialized, and very few samples are collected. Once a positive reward appears, the optimization direction will be absolutely overwhelmingly in that direction, resulting in an increasing number of biased samples in the sample pool, forming a negative feedback loop. In fact, the sample layer is estimating expectations in the form of a Monte Carlo method. Therefore, the early stages should be more random and exploratory in sample collection. Therefore, we designed the early actions to be generated in the form of a uniform distribution. That is, the agent's action has a probability p of being generated from a uniform distribution, and the probability of taking the action with the maximum Q value is 1-p. This p is designed to gradually decrease as the training step increases. When the step reaches a certain threshold (model hyperparameter), it is reset to zero, achieving a smooth adaptive effect.
[0085] Because all samples are generated dynamically online, the model can easily fall into a localized trap. This occurs when biased samples are trained, creating biased samples again, and training continues in this order. To overcome this trap, we designed a ReplayBuffer mechanism inspired by DQN. This mechanism randomly samples from a larger step range, mitigating training anomalies caused by localized sample bias. Of course, after each interaction step, we need to store the corresponding episode data in a buffer for subsequent sampling.
[0086] Because this interaction runs through the entire engine pipeline, its performance is crucial. We leveraged parallel computing between agents, aggregating episode data across compute nodes and writing it to a shared buffer. This significantly improved computing performance across multiple agents, a key advantage of the distributed prediction architecture.
[0087] In a preferred embodiment of the present application, as shown in Figure 4, the model training layer maps elements of the reinforcement learning framework, such as agent states and observation information, to a vector space, uses the Qmix network for learning, and designs a padding algorithm to solve dynamic agent training. Figure 4 shows the design of the qmixer module (i.e., Qmix network). The structure described in (b) is the structure of the entire module. Figure (a) shows the mixer subnetwork, which is used to fuse the Q values of each agent and merge the Q values of all agents into a single Q value (scalar). Figure (c) shows the modeling of a single agent subnetwork.
[0088] Since our problem involves learning across multiple agents, the goal of each agent in classic RL (reinforcement learning) problems is to optimize its Q-value, maximizing its own Q-value through iterative optimization of policy gradients. However, this problem becomes much more complex when applied to multiple agents. The question arises about how to formulate an optimization goal. Currently, the most mainstream solution in the industry is a centralized training and distributed usage architecture. This architecture offers the best generalization capabilities, the highest engineering efficiency, and avoids the complexities of communication between different agents.
[0089] Therefore, our algorithm is also designed based on this architecture. We define a global optimization goal, Qtot, which is the Q value that needs to be satisfied:
[0090] That is, the Q function of each agent is monotonically increasing. This ensures that when training and predicting, we only need to focus on optimizing the Q value of each agent to optimize the global goal. The optimization relationship is described as follows:
[0091] For the design of the mixer sub-network, our goal is to design a fusion network, input a series of agent Q values, and obtain an overall Q value through a function transformation. This function can of course be directly fitted using a NN. Here, a two-layer MLP network is designed, and the gradient is used to ensure the monotonically increasing nature of the function.
[0092] The design of the Agent subnetwork mainly inputs an observation tensor and finally outputs a predicted Q value for each action in the action set. The policy gradient of the Q function is used to optimize and maximize the Q value. This can be completely abstracted as a classification subnetwork. Therefore, the algorithm design is also based on the classification structure, using an MLP network plus a logits module. However, this is actually a sequence problem. Before the model forms the current action, it must also obtain information about past actions. Therefore, the network also adds a GRU network module to capture information about historical actions.
[0093] In order to enable the model to support a dynamic number of agents, we designed a padding algorithm to solve the problem of dynamic agent training.
[0094] Specifically
[0095] First, set a maximum number of agents N supported. This number is a hyperparameter, so an optimal value can be tried out in experiments. Each time a sample with n layers is input, only the first n of the N agents are activated, and the rest are inactive. To achieve this inactive state, it is necessary to ensure that any computational output of these inactive agents cannot affect the training prediction of the model. The simplest implementation idea is to set the gradients generated by these agents to zero. We designed a masking mechanism to cleverly achieve this through the characteristics of automatic differentiation of the computational graph. The schematic diagram of the masking mechanism is shown in Figure 5.
[0096] As shown in Figure 5, four tensors are described, where QTensor is the tensor formed by the Q values of all agents, mask is a tensor containing only 0 and 1, output is the output tensor of the Hadamard product operation of the two, and QGrad is the differential tensor of QTensor.
[0097] Arrows marked with squares depict forward computations, while arrows marked with circles depict the equivalent automatic differentiation process. In the computational graph, to prevent high-order Jacobians from appearing in differential operations between tensors, the computational graph automatically converts all tensor-to-tensor differentials into scalar-to-tensor differentials. This results in all tensor differentials having the same shape as the tensor itself. Converting a tensor to a scalar is essentially performing a reduceSum operation on it using a tensor of all 1s. Finally, the arrow marked with a triangle represents the derivative of the scalar with respect to the original tensor, resulting in the final differential tensor.
[0098] As can be seen from the figure, the mask tensor we designed participates in the calculation of the gradient, which is equivalent to implementing a layer of Hook. In this way, we can selectively set some gradients to zero by controlling the distribution of 0 in the mask, thereby shielding the inactive agent at the gradient level.
[0099] For further information, see Figure 6. The QLearner module is the core training component of the entire algorithm engine and utilizes the currently mainstream DQN training architecture. It trains and optimizes the agent subnetwork to maximize the Q value for the optimal policy action. The Q network and Target network in the figure are both encapsulations of the agent subnetwork and share the same structure, except that the Target network precedes the Q network by one time step to estimate the Q value at time T+1. The Environment module definition is consistent with that used in RL problems. It takes an action as input and generates the state, observation, and corresponding reward resulting from the current action. This module is essentially a sequence generator, continuously generating samples and buffering them for DQN training. DQN uses a time-determined learning approach, with residual error as the final loss function.
[0100] The design of reward directly affects the convergence of the model. The previous section explains the design ideas of reward. The following section gives the specific calculation implementation of reward: d=D cur -D prev
[0101] Here, Ri is the reward of the i-th agent. As previously demonstrated, we only need to focus on the rewards of each agent during training. Therefore, when designing the reward module, we also follow the distribution plus final summary approach. The final accumulation of Ri is the reward of the entire system.
[0102] The reward is calculated using a progressive incentive mechanism, which means that the reward can be positive or negative at any time. That is, at any time, the distance between the agent's current position and the target position is calculated and recorded in the cache. At the same time, it is compared with the distance value recorded last time in the cache. If the distance is smaller than the previous distance, it is considered that a positive progressive incentive should be given, and a reward of size d is given. Of course, this d value will be normalized in the implementation to ensure the stability of the numerical calculation. Conversely, when d is negative, it is considered that the current agent has gone in the wrong direction, and corresponding negative feedback will be given. However, this negative feedback is very dense in the early stage of model training, and the positive rewards in the early stage are also very sparse. If no special treatment is done, the model will quickly fall into an oscillation state in the early stage. Therefore, an adaptive attenuation weight k is made here. This value will be small in the early stage of training and will gradually increase in the later stage of training, reaching a maximum of 1. This not only plays the role of negative reward punishment, but also ensures the stability of early training.
[0103] Please refer to Figure 7 for the overall implementation of the model service module. It is used to serve users with models trained by the trainer through a RESTful interface. First, our trained model is packaged into a mar package along with some data pipeline hooks. If it is the first time it is launched online, the model will register its information with the engine's Register module so that the model service can detect that a new model needs to be served. The model package is then uploaded to HDFS. Next, we implemented a version control module. Each model launched online has a unique version number. The version module provides functions such as version hot switching and version rollback, ensuring seamless online model switching and zero-cost fault-tolerant rollback. By default, the latest version of the model is used for serving. The monitor is a background daemon responsible for polling the version module. Once a new version is uploaded, it immediately downloads the model to the local disk of the service node. At this time, the front-end RESTful service interface can load the model and its calculation code into memory to provide model calculation services for user requests. In addition, the service can specify the version to access a specific version of the model in the version module and can also initiate an emergency rollback request to the version module.
[0104] Referring to FIG. 2 , the present application also provides a banner synthesis device based on reinforcement learning, comprising:
[0105] A preprocessing module is used to perform layer preprocessing on the sample and expand the sample, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping;
[0106] The feature module is used to extract multimodal features from the expanded samples and represent the multimodal features with feature vectors. The multimodal features include structural features, design text features, and visual features.
[0107] The sample enhancement module is used to solve the biased training problem of multimodal samples and generate reinforcement learning training samples online using dynamic sampling.
[0108] Model training module, used to train the model and generate an optimized model;
[0109] The service module provides services to users through the Restful interface after the optimized model.
[0110] The above is merely a preferred embodiment of the present application. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present application, and such improvements and modifications should be considered within the scope of protection of the present application. Structures, devices, and operating methods not specifically described or explained in this application shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.
[0111] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0112] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A banner synthesis method based on reinforcement learning, characterized in that: include: S1, preprocessing the sample through layers and expanding the sample, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping; S2. Extract multimodal features from the expanded samples through the encoder, represent the multimodal features with feature vectors, and enhance the generalization ability of the model through multimodal feature learning. The multimodal features include structural features, design text features, and visual features. S3, solve the biased training problem for the samples extracted from multimodal, and use dynamic sampling to generate samples for reinforcement learning training online; S4, establish a model training layer, train the training model, and generate an optimized model; S5. Provide services to users through the optimized model in the form of a Restful interface.
2. The banner synthesis method based on reinforcement learning according to claim 1, characterized in that: The preprocessing steps of layer alignment, merging and cropping from the sample include: Get an approximate KV, filter the artboards, determine which artboard is the main KV, and align the layers; Each layer is scaled based on the size ratio between the main KV and the artboard to be processed.
3. The banner synthesis method based on reinforcement learning according to claim 2, characterized in that: The step of performing sample expansion comprises: Obtain available main KVs from the sample so that each main KV has several label data of different sizes on average; The above main KVs are paired in a pairwise manner to generate samples, that is, to obtain enhanced sample data.
4. The banner synthesis method based on reinforcement learning according to claim 1, characterized in that: The encoder adopts a hybrid DNN network design feature to encode the extracted multimodal features in a DNN manner, wherein the multimodal features are represented by feature vectors and the multimodal features need to be mapped to a unified vector space.
5. The banner synthesis method based on reinforcement learning according to claim 1, characterized in that: The biased training described above adopts a replayBuffer mechanism, and the dynamic sampling method adopts a random sampling method that combines uniform sampling and optimized sampling.
6. The banner synthesis method based on reinforcement learning according to claim 1, characterized in that: The model training layer maps the elements of the reinforcement learning framework such as the agent's state and observation information to the vector space, uses the Qmix network for learning, and designs a padding algorithm to solve the problem of dynamic agent training.
7. The banner synthesis method based on reinforcement learning according to claim 6, characterized in that: The Qmix network is used for learning using a centralized training distributed usage architecture, which specifically includes the following steps: Define a global optimization goal, Qtot, where the Q value needs to satisfy: That is, the Q function Qa of each agent is guaranteed to be monotonically increasing; Optimize the Q value of each agent. The optimization relationship is as follows: Design the mixer subnetwork to obtain the overall Q value and ensure the monotonically increasing function; Design the Agent sub-network to optimize and maximize the Q value; A GRU network module is also added to the Agent sub-network to capture historical action information.
8. The banner synthesis method based on reinforcement learning according to claim 7, characterized in that: The steps of designing the mixer sub-network include: Design a fusion network, input a series of optimized Q values, and obtain an overall Q value through a function transformation; Design a two-layer MLP network and use the gradient to ensure the monotonically increasing function.
9. The banner synthesis method based on reinforcement learning according to claim 6, characterized in that: The method of designing a padding algorithm to solve the problem of dynamic agent training includes: First, set a maximum number of agents N that can be supported. This number is a hyperparameter, so you can try to find the best value in the experiment. Each time you input a sample with n layers, only the first n of the N agents are activated, and the rest are inactive. The QLearner module is used to train and optimize the agent subnetwork to maximize the Q value of the optimal strategy action. Input action through the Environment module to obtain the state, observation and corresponding reward caused by the current action.
10. A banner synthesis device based on reinforcement learning, characterized in that: The banner synthesis method based on reinforcement learning applied to any one of claims 1 to 9 above, the banner synthesis device based on reinforcement learning comprises: A preprocessing module is used to perform layer preprocessing on the sample and expand the sample, wherein the preprocessing includes initial offset calculation, agent size change, layer merging, and cropping; The feature module is used to extract multimodal features from the expanded samples and represent the multimodal features with feature vectors, where the multimodal features include structural features, design text features, and visual features; The sample enhancement module is used to solve the biased training problem of the samples extracted from the multimodal model and to generate samples for reinforcement learning training online using dynamic sampling. The model training module is used to train the model and generate an optimized model; The service module provides services to users through the Restful interface using the optimized model.
Citation Information
Patent Citations
Picture generation method and system, electronic equipment and storage medium
CN110750666A
Poster generation model training method, poster generation method and related equipment
CN115049043A
Banner synthesis method and device based on reinforcement learning
CN117876810A
Generating personalized banner images using machine learning
US20210103957A1