Picture and video-oriented advertisement material batch generation system and method

By using semantic enhancement processing and iterative memory modules to deeply analyze advertising creative generation requirements, the problem of inconsistent batch generation of advertising creatives in existing technologies has been solved, enabling high-quality multi-platform, multi-size, and multi-format creative generation.

CN122156373APending Publication Date: 2026-06-05ONEWAY NETWORK(HK) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ONEWAY NETWORK(HK) LTD
Filing Date
2026-02-25
Publication Date
2026-06-05

Smart Images

  • Figure CN122156373A_ABST
    Figure CN122156373A_ABST
Patent Text Reader

Abstract

The application discloses a kind of picture and video-oriented advertising material batch generation system and method, mainly related to advertising design development technical field.The system includes: semantic enhancement processing module, for obtaining advertising material generation demand, executing explicit and implicit semantic enhancement processing, determining advertising material generation demand semantic feature set;Picture batch iteration memory processing module is used for executing batch iteration memory processing, and obtaining picture advertising material batch generation result;Video batch iteration memory processing module is used for obtaining video advertising material batch generation result;Material consistency detection module is used for carrying out material consistency detection, when material consistency detection passes, obtains target advertising material batch generation result.The beneficial effects of the present application are that the technical problems that the quality of advertising material batch generation is not the same in the prior art, cannot fully meet the demand, and the technical effect of improving the quality of advertising material batch generation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of advertising design and development technology, specifically to a system and method for batch generation of advertising materials for images and videos. Background Technology

[0002] With the rapid development of the internet advertising industry, advertising design and development faces the demand for advertising materials across multiple platforms, sizes, and formats. The batch generation of image and video advertising materials has become a key technological aspect. Currently, mainstream methods for batch generating advertising materials typically employ template-based rendering and splicing techniques. These methods use pre-set image and video editing templates to automatically fill and combine elements such as product images, copy, and brand logos, enabling the production of a large volume of advertising materials in a short period.

[0003] However, current methods primarily focus on analyzing explicit parameters of advertising needs, such as structured information like size, resolution, file format, and brand logo placement. The resulting creative materials often fail to adequately align with advertising requirements, and the batch-generated materials are frequently fragmented, leading to inconsistent use and unsatisfactory results. Consequently, the problem of limited usable batch-generated materials often arises, lowering the overall quality of bulk advertising creative generation. Summary of the Invention

[0004] This application provides a system and method for batch generation of advertising creatives for images and videos, which addresses the technical problem that the quality of batch-generated advertising creatives in the prior art is inconsistent and cannot fully meet the needs.

[0005] In view of the above problems, this application provides a system and method for batch generation of advertising creatives for images and videos.

[0006] The first aspect of this application provides a system for batch generation of advertising creatives for images and videos, the system comprising: The semantic enhancement processing module is used to obtain the advertising material generation requirements, perform explicit and implicit semantic enhancement processing, and determine the semantic feature set of the advertising material generation requirements. The image batch iterative memory processing module is used to perform batch iterative memory processing on the image advertising material generation task based on the semantic feature set of the advertising material generation requirements to obtain the batch generation results of image advertising materials. The video batch iterative memory processing module is used to perform batch iterative memory processing on the video advertising material generation task based on the semantic feature set of the advertising material generation requirements to obtain the batch generation results of video advertising materials. The material consistency detection module is used to traverse the batch generation results of image advertising materials and the batch generation results of video advertising materials and perform material consistency detection on each. When the material consistency detection passes, the batch generation results of image advertising materials and the batch generation results of video advertising materials are summarized to obtain the batch generation results of the target advertising materials.

[0007] A second aspect of this application provides a method for batch generation of advertising creatives for images and videos, the method comprising: The process involves: acquiring advertising creative generation requirements, performing explicit and implicit semantic enhancement processing to determine the semantic feature set of advertising creative generation requirements; based on the semantic feature set of advertising creative generation requirements, performing batch iterative key memory processing on the image advertising creative generation task to obtain batch generation results of image advertising creatives; based on the semantic feature set of advertising creative generation requirements, performing batch iterative key memory processing on the video advertising creative generation task to obtain batch generation results of video advertising creatives; traversing the batch generation results of image advertising creatives and video advertising creatives respectively to perform creative consistency checks; when the creative consistency check passes, summarizing the batch generation results of image advertising creatives and video advertising creatives to obtain the batch generation results of the target advertising creatives.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: The system provided in this application includes a semantic enhancement processing module for obtaining advertising material generation requirements, performing explicit and implicit semantic enhancement processing, and determining the semantic feature set of advertising material generation requirements; an image batch iterative memory processing module for performing batch iterative memory processing on image advertising material generation tasks based on the semantic feature set of advertising material generation requirements, and obtaining batch generation results of image advertising materials; a video batch iterative memory processing module for performing batch iterative memory processing on video advertising material generation tasks based on the semantic feature set of advertising material generation requirements, and obtaining batch generation results of video advertising materials; and a material consistency detection module for iterating through the batch generation results of image advertising materials and video advertising materials and performing material consistency detection on each. When the material consistency detection passes, the batch generation results of image advertising materials and video advertising materials are aggregated to obtain the batch generation results of the target advertising materials. This achieves the technical effect of deeply analyzing advertising material generation requirements from both explicit and implicit dimensions, improving the coherence of batch-generated materials, and improving the quality of material generation. Attached Figure Description

[0009] Appendix Figure 1 This is a schematic diagram of the structure of a batch generation system for advertising materials for images and videos provided in an embodiment of the present invention.

[0010] Appendix Figure 2 This is a schematic diagram of a method for batch generating advertising materials for images and videos provided in an embodiment of the present invention.

[0011] The labels shown in the attached diagram: Semantic enhancement processing module 11, image batch iterative memory processing module 12, video batch iterative memory processing module 13, and material consistency detection module 14. Detailed Implementation

[0012] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims. It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices.

[0013] Example 1, as shown in the appendix Figure 1 As shown, this application provides a system for batch generation of advertising creatives for images and videos, the system comprising: Semantic enhancement processing module 11 is used to obtain advertising material generation requirements, perform explicit and implicit semantic enhancement processing, and determine the semantic feature set of advertising material generation requirements; Image batch iterative memory processing module 12 is used to generate a set of semantic features of demand based on the advertising material, perform batch iterative memory processing on the image advertising material generation task, and obtain the batch generation result of image advertising material; The video batch iterative memory processing module 13 is used to generate a set of semantic features of demand based on the advertising material, perform batch iterative memory processing on the video advertising material generation task, and obtain the batch generation result of the video advertising material. The material consistency detection module 14 is used to traverse the batch generation results of the image advertising materials and the batch generation results of the video advertising materials and perform material consistency detection respectively. When the material consistency detection passes, the batch generation results of the image advertising materials and the batch generation results of the video advertising materials are summarized to obtain the batch generation results of the target advertising materials.

[0014] Furthermore, the semantic enhancement processing module 11 also includes: a formatting processing unit, used to uniformly format the natural language creative briefing data, structured advertising configuration data, and platform specification data in the advertising material generation requirements to construct a raw requirement data set; an explicit semantic enhancement processing unit, used to perform explicit semantic enhancement processing on the raw requirement data set to obtain a first requirement semantic feature set, wherein the first requirement semantic feature set includes visual rule features, layout matching features, and style control features; an implicit semantic enhancement processing unit, used to perform implicit semantic enhancement processing on the raw requirement data set to construct a second requirement semantic feature set; and a union calculation unit, used to perform union calculation on the first requirement semantic feature set and the second requirement semantic feature set to determine the advertising material generation requirement semantic feature set.

[0015] Furthermore, the implicit semantic enhancement processing unit also includes: a retrieval subunit, used to retrieve the historical advertising material repository using the original demand data set as an index to obtain a historical advertising material set and a campaign performance data set; a filtering subunit, used to filter the campaign performance data set according to preset campaign performance constraints to obtain a valid campaign performance data set; a mapping extraction subunit, used to perform mapping extraction on the historical advertising material set based on the valid campaign performance data set to obtain a valid historical advertising material set; a feature extraction subunit, used to perform implicit semantic feature extraction on the valid historical advertising material set to obtain an implicit semantic feature set; and a semantic feature set acquisition subunit, used to perform implicit semantic enhancement processing on the implicit semantic feature set to obtain the second demand semantic feature set.

[0016] Furthermore, the semantic feature set acquisition subunit also includes: a feature mean processing microunit, used to traverse the implicit semantic feature set and perform feature mean processing to determine the initial implicit semantic enhancement center; an implicit semantic feature extraction microunit, used to extract implicit semantic features from the implicit semantic feature set whose Euclidean distance to the initial implicit semantic enhancement center satisfies a preset nearest neighbor bandwidth, and construct the neighborhood of the initial implicit semantic enhancement center; and an iterative implicit semantic enhancement center neighborhood construction microunit, used to take the implicit semantic features in the neighborhood of the initial implicit semantic enhancement center corresponding to the maximum Euclidean distance to the initial implicit semantic enhancement center as the iterative implicit semantic feature... A semantic enhancement center is established, and an iterative implicit semantic enhancement center neighborhood is constructed. A stopping condition determination micro-unit is used to compare the size of the neighborhood features of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood to determine whether the preset enhancement processing stopping condition is met. If it is met, the initial implicit semantic enhancement center neighborhood is used as the second required semantic feature set. If it is not met, the implicit semantic features corresponding to the maximum Euclidean distance between the iterative implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center are used to continue the enhancement processing of the iterative implicit semantic enhancement center until the preset enhancement processing stopping condition is met.

[0017] Furthermore, when the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is greater than or equal to the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, the preset enhancement processing stop condition is met; when the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is less than or equal to a preset difference threshold, the preset enhancement processing stop condition is met; when the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is greater than a preset difference threshold, the preset enhancement processing stop condition is not met.

[0018] Furthermore, the explicit semantic enhancement processing unit also includes: a keyword extraction subunit, used to extract keywords from the original set of requirements to obtain an explicit requirement keyword set; a constraint filtering subunit, used to perform same-dimensional constraint filtering on the explicit requirement keyword set from three dimensions: visual, layout, and style, to obtain a filtered explicit requirement keyword set; and a structured encoding processing subunit, used to perform structured encoding processing on the filtered explicit requirement keyword set to obtain the first requirement semantic feature set.

[0019] Furthermore, the image batch iterative memory processing module 12 also includes: an image advertising element set construction unit, used to construct an image advertising element set based on the semantic feature set of the advertising material generation requirements; a first-order semantic parsing unit, used to generate a first image advertising material according to the image advertising element set, perform first-order semantic parsing on the first image advertising material to obtain the semantic features of the first image advertising material, add the semantic features of the first image advertising material into an initially empty vector, and construct a first semantic memory unit; a second image advertising material generation unit, used to generate a second image advertising material according to the first semantic memory unit and the image advertising element set; an iterative update unit, used to iteratively update the first semantic memory unit based on the second image advertising material to obtain a second semantic memory unit; and an image advertising material batch generation result acquisition unit, used to generate a third image advertising material according to the second semantic memory unit and the image advertising element set, and so on, performing material generation and semantic memory unit iterative update until the preset batch generation quantity of material in the image advertising material generation task is met, and obtaining the image advertising material batch generation result.

[0020] Furthermore, the second image advertising material generation unit also includes: a first image advertising material semantic feature extraction subunit, used to extract the semantic features of the first image advertising material from the first semantic memory unit; a weight assignment subunit, used to assign associated weights to the semantic features of the first image advertising material and the set of image advertising elements to obtain an element association weight set; and a material generation subunit, used to generate the second image advertising material based on the element association weight set and the set of image advertising elements.

[0021] Furthermore, the material consistency detection module 14 also includes: a multi-dimensional feature extraction unit, used to traverse the batch generation results of image ad materials and video ad materials to perform multi-dimensional feature extraction, and obtain multi-dimensional features of image ad materials and video ad materials; a similarity calculation unit, used to calculate the similarity between the multi-dimensional features of image ad materials and video ad materials, and obtain multi-dimensional feature similarity; a pass detection unit, used to ensure that the material consistency detection passes when the multi-dimensional feature similarity is greater than or equal to a preset similarity threshold; and a fail detection unit, used to ensure that the material consistency detection fails when the multi-dimensional feature similarity is less than the preset similarity threshold.

[0022] Example 2, based on the same inventive concept as the batch generation system for advertising materials for images and videos in the foregoing examples, as shown in the appendix. Figure 2 As shown, this application provides a method for batch generation of advertising materials for images and videos. The method and system embodiments in this application are based on the same inventive concept. The method includes: Obtain the advertising creative generation requirements, perform explicit and implicit semantic enhancement processing, and determine the semantic feature set of the advertising creative generation requirements; Furthermore, the ad creative generation requirements are obtained, and explicit and implicit semantic enhancement processing is performed to determine the set of semantic features of the ad creative generation requirements, including: The natural language creative briefing data, structured ad configuration data, and platform specification data in the advertising material generation requirements are uniformly formatted to construct the original data set of the requirements; The original data set of requirements is subjected to explicit semantic enhancement processing to obtain a first set of semantic features of requirements, wherein the first set of semantic features of requirements includes visual rule features, layout matching features and style control features. The original data set of requirements is subjected to implicit semantic enhancement processing to construct a second set of semantic features of requirements; The first set of semantic features of demand and the second set of semantic features of demand are combined to determine the set of semantic features of demand for generating advertising materials.

[0023] In one embodiment, ad creative generation requirements refer to the raw information input by operations personnel to guide ad creative generation, including natural language creative briefing data, structured ad configuration data, and platform specification data. Natural language creative briefing data, such as data with a strong technological feel and suitability for young people, is used to describe the requirements. Structured ad configuration data specifies the platform configurations for image and video ads, such as platform type, size specifications, and brand elements. The platform specification data outlines the platform's requirements for ads, such as resolution and file size limits. The semantic feature set of ad creative generation requirements reflects customer needs from multiple dimensions.

[0024] In one embodiment, because the sources of natural language creative briefing data, structured ad configuration data, and platform specification data in the ad creative generation requirements are different, a unified formatting process is needed to convert them into a parsable, similar format for easier subsequent in-depth analysis of the requirements. Preferably, taking energy drink ad requirements as an example, the natural language creative briefing data is for an energy drink that makes programmers working late feel saved, with a technological feel but not too cold. This is input into a BERT model accessible to those skilled in the art, and after sequence labeling and text classification, a structured tag set is output: Scene: Late night office, Target profession: Programmer, Emotional starting point: Fatigue, Emotional ending point: Awakening, Product category: Energy drink, Style theme: Technological, Style constraint: Avoid coldness. The structured ad configuration data is input in JSON format, such as the platform selection, budget amount, start and end dates, target audience profile tags, etc. The platform specification data is an API document, including aspect ratio, resolution threshold, file format whitelist, duration range, bitrate limit, file size limit, etc. After field mapping and unit normalization, these are merged and stored in the original data set of the requirements.

[0025] The original demand data set can be intuitively analyzed to identify the demand for advertising creatives. However, due to the diverse sources, overlapping or conflicting parts may appear in the data. Explicit semantic enhancement processing is needed to reduce data redundancy and achieve the goal of providing accurate guidance for the subsequent batch generation of advertising creatives. Preferably, the analysis is performed from three dimensions: visual, layout, and style, to obtain a first set of semantic features for demand. This first set of semantic features includes visual rule features, layout matching features, and style control features.

[0026] Meanwhile, the original demand data set also implicitly contains some information that needs further analysis to guide the generation of advertising materials. By performing implicit semantic enhancement processing, information such as platform preferences and high conversion bias of advertisements implied in historical advertising data and campaign performance can be mined to obtain the second demand semantic feature set.

[0027] Furthermore, the semantic feature set of the first requirement and the semantic feature set of the second requirement are combined to obtain a complete semantic feature set of advertising material generation requirements. This enables in-depth analysis of advertising requirements from both explicit and implicit dimensions, namely, intuitive and implicit information mining. This reduces the increased cost of repeated generation due to mismatch between generated materials and requirements, and achieves the technical effect of improving the reliability of batch generation of advertising materials.

[0028] Furthermore, implicit semantic enhancement processing is performed on the original set of demand data to construct a second set of demand semantic features, including: Using the original data set of the requirements as an index, the historical advertising material repository is retrieved to obtain the historical advertising material set and the set of campaign performance data; The effective campaign performance data set is obtained by filtering the data set according to the preset campaign performance constraints. Based on the effective delivery performance data set, the historical ad creative set is mapped and extracted to obtain the effective historical ad creative set; Implicit semantic features are extracted from the set of valid historical advertising materials to obtain an implicit semantic feature set; The implicit semantic feature set is subjected to implicit semantic enhancement processing to obtain the second required semantic feature set.

[0029] It's important to note that the historical ad creative repository is a database used to store past ad images, video creatives, and their corresponding delivery data. Each creative is linked to and stored with performance metrics such as click-through rate (CTR), conversion rate, and completion rate. Preset performance constraints refer to threshold rules used to select high-quality samples; these can be optional, such as CTR ≥ 3%, conversion rate ≥ 2%, and completion rate ≥ 40%. Implicit semantic features are visual, layout, or rhythmic features extracted from effective historical ad creatives that are not explicitly expressed in the current requirements but significantly impact performance. Examples include color proportion features, layout proportion features, shot rhythm features, and text emotional intensity features.

[0030] Using industry tags, platform types, and style keywords from the original demand data set as search criteria, a structured conditional search is performed in the historical ad creative repository to obtain a set of matching historical ad creatives and their corresponding performance data. For example, 500 3C+tech-themed creatives from the past year are selected from the historical ad creative repository, and their click-through rate (CTR), conversion rate (CTR), and completion rate (CTR) data are extracted simultaneously. Then, the data is filtered according to preset performance constraints, such as setting a CTR ≥3%, a conversion rate ≥1.5%, and a CTR ≥35% as valid criteria, resulting in a valid performance data set with 120 valid performance data points.

[0031] Furthermore, based on the effective campaign performance data set, a reverse index is performed on the historical ad creative set to obtain the corresponding effective historical ad creative set. Then, by extracting implicit information from the creatives in the effective historical ad creative set, an implicit semantic feature set is obtained.

[0032] Preferably, for image materials, their compositional features, color distribution, visual saliency heatmaps, OCR text regions, faces, and gaze directions are extracted; for video materials, their shot boundaries, keyframes, motion vectors, audio energy, and subtitle timelines are extracted. For example, the material is a WeChat Moments ad image depicting a programmer working late at night. The extracted visual saliency heatmap shows the highest concentration of attention in the lower right third of the image, corresponding to the product location. OCR identifies the text containing the specific time point 3:17 AM, and face detection shows the person's gaze is pointing towards the lower right corner. These heterogeneous features are uniformly encoded into 256-dimensional embedding vectors as implicit semantic features.

[0033] Since the implicit semantic feature set includes various scenarios of valid historical ad creative sets that meet the original data set requirements, including some accidental cases or errors, this set needs to be filtered, which is to perform implicit semantic enhancement processing, in order to obtain the second set of semantic features required. Thus, the technical effect of filtering implicit information, mining statistically significant implicit semantic features, assisting in the subsequent generation of ad creatives, and improving the quality of generated creatives is achieved.

[0034] Furthermore, implicit semantic enhancement processing is performed on the implicit semantic feature set to obtain the second requirement semantic feature set, including: The implicit semantic feature set is traversed and the feature mean is processed to determine the initial implicit semantic enhancement center; Extract implicit semantic features from the implicit semantic feature set whose Euclidean distance to the initial implicit semantic enhancement center satisfies the preset nearest neighbor bandwidth, and construct the neighborhood of the initial implicit semantic enhancement center; The implicit semantic features corresponding to the maximum Euclidean distance from the initial implicit semantic enhancement center in the neighborhood of the initial implicit semantic enhancement center are used as iterative implicit semantic enhancement centers, and the neighborhood of the iterative implicit semantic enhancement center is constructed. Compare the neighborhood feature values ​​of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood to determine whether the preset enhancement processing stop condition is met. If it is met, the initial implicit semantic enhancement center neighborhood is used as the second required semantic feature set. If the condition is not met, the implicit semantic features corresponding to the maximum Euclidean distance between the iterative implicit semantic enhancement center and the neighborhood of the iterative implicit semantic enhancement center are used to continue the enhancement process for the iterative implicit semantic enhancement center until the preset enhancement process stop condition is met.

[0035] Furthermore, when the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is greater than or equal to the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, the preset enhancement processing stop condition is met. When the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is less than or equal to a preset difference threshold, the preset enhancement processing stop condition is met. When the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is greater than a preset difference threshold, the preset enhancement processing stop condition is not met.

[0036] In one embodiment of this application, it should be noted that feature mean processing refers to averaging all implicit semantic features along their dimensions to obtain a central vector. In other words, the initial implicit semantic enhancement center is used to represent the average semantic trend of the overall samples. Euclidean distance is used to measure the degree of difference between two feature vectors; the smaller the distance, the more semantically similar they are. Preset nearest neighbor bandwidth refers to a threshold parameter used to define the range of the central neighborhood. For example, samples with a distance less than 0.8 are set as nearest neighbor samples and added to the neighborhood. The neighborhood feature quantity can be understood as the number of implicit semantic features contained in the neighborhood. The preset difference threshold is a configurable hyperparameter, and its setting needs to strike a balance between pattern purity and computational efficiency. It should be set by those skilled in the art. A smaller threshold results in more rigorous algorithm convergence, more iterations, and higher consistency of the extracted core pattern features; a larger threshold results in faster convergence and lower computational overhead, but may prematurely stop in an incompletely converged state.

[0037] Preferably, the implicit semantic feature set is traversed and its dimensions are arithmetically averaged to generate an initial implicit semantic enhancement center. This center represents the average creative profile of currently retrieved historical successful cases in the semantic space, i.e., the common values ​​of these cases in dimensions such as visual style, emotional intensity, and symbol density. Based on a preset nearest neighbor bandwidth, the Euclidean distance between the initial implicit semantic enhancement center and each implicit semantic feature in the implicit semantic feature set is calculated, and all features with a distance less than or equal to this bandwidth are extracted to form the neighborhood of the initial implicit semantic enhancement center.

[0038] The feature vector with the largest Euclidean distance from the initial center within the initial neighborhood is identified and set as the first iterative implicit semantic enhancement center. This first iterative implicit semantic enhancement center, i.e., the edge feature, represents the successful case in the current neighborhood that differs the most from the average experience but is still included in the similarity range. Based on the same principle as constructing the neighborhood of the initial implicit semantic enhancement center, the neighborhood of the iterative implicit semantic enhancement center is constructed again.

[0039] By moving the center point toward the neighborhood boundary, we are essentially performing a drift operation toward a high-density region. If there is a denser sample distribution in the edge direction, the new neighborhood will cover more feature vectors; if not, we can also perform exclusion verification.

[0040] The algorithm compares the neighborhood feature values ​​of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood, and determines whether to terminate the iteration based on a preset enhancement processing stop condition. The stop condition includes three scenarios: First, when the initial neighborhood feature value is greater than or equal to the iterative neighborhood feature value, it indicates that the current movement direction has crossed the density peak region, and continued drifting will lead to neighborhood shrinkage, thus satisfying the preset enhancement processing stop condition. In this case, the iteration is terminated, and the initial implicit semantic enhancement center neighborhood is output as the second required semantic feature set. For example, suppose the initial neighborhood feature value is 72 and the iterative neighborhood feature value is 68, where 68 is less than 72. It is determined that the density peak has been crossed, and continuing to move in the current direction will cause neighborhood shrinkage, meaning the previous center is closer to the true density peak. In this case, the iteration is immediately terminated, and the implicit semantic features included in the initial implicit semantic enhancement center neighborhood are output as the second required semantic feature set. This achieves the technical effect of avoiding continued drifting to low-density areas after crossing the peak, and extracting marginalized and unrepresentative features.

[0041] Second, when the initial neighborhood feature quantity is less than the iterative neighborhood feature quantity, and the difference between the two is less than or equal to a preset difference threshold set by those skilled in the art, it indicates that the neighborhood is still expanding but the marginal gain is below an acceptable level. Continuing the iteration will have limited improvement on feature purity and will generate additional computational overhead. At this time, the iteration is terminated and the initial implicit semantic enhancement center neighborhood is output as the second required semantic feature set. Third, when the initial neighborhood feature quantity is less than the iterative neighborhood feature quantity, and the difference between the two is greater than a preset difference threshold, it indicates that the neighborhood is still expanding significantly and the current center has not yet reached the density peak region. The system uses the implicit semantic feature farthest from the center in the current iterative neighborhood as the new iterative center and repeats the neighborhood construction and feature quantity comparison until one of the first two stopping conditions is met.

[0042] When the preset enhancement processing stopping condition is met, the system outputs the set of implicit semantic feature vectors covered by the current neighborhood as the second set of required semantic features. This set is a subset obtained after density shift convergence of the original implicit semantic feature set, with higher sample density, stronger feature consistency, and effective removal of marginal outliers. This avoids the influence of extreme outliers; a single high-performing but stylistically diverse successful case can significantly skew the average center. Through implicit semantic enhancement processing, representative features are selected as the second set of required semantic features, achieving the technical effect of providing high-quality required semantics for subsequent material generation.

[0043] Furthermore, the original set of demand data undergoes explicit semantic enhancement processing to obtain a first set of demand semantic features, including: Keyword extraction is performed on the original set of requirements to obtain a set of explicit requirement keywords; From the three dimensions of visual, layout and style, the explicit requirement keyword set is traversed and constrained by the same dimension to obtain the filtered explicit requirement keyword set. The set of explicit requirement keywords is subjected to structured encoding to obtain the first set of requirement semantic features.

[0044] Preferably, the natural language creative brief in the original demand data set is first segmented and keywords are extracted. For example, if the input text is "create a high-end tech-style poster targeting young users," the BERT keyword extraction model or TF-IDF algorithm is used to identify a set of keywords, such as "high-end," "tech-style," "poster," and "young users." Then, this explicit demand keyword set is mapped to three preset semantic dimensions: visual dimension (e.g., tech-style corresponds to cool colors and metallic texture), layout dimension (e.g., poster corresponds to a single-page structure and centered title), and style dimension (e.g., high-end corresponds to low saturation and a white space ratio of ≥40%).

[0045] Preferably, the explicit requirement keyword set is divided along the same dimension according to visual, layout, and style to obtain a visual explicit requirement keyword set, a layout explicit requirement keyword set, and a style explicit requirement keyword set. Then, each set undergoes constraint-based filtering within the same dimension. For example, if the visual dimension includes both cool and warm colors, features with higher frequency of occurrence are retained, while conflicting features are removed. After filtering, the keywords are converted into structured parameters, such as cool color ratio = 0.65, white space ratio = 0.45, title position = Center, style level = LuxuryLevel2, etc., forming the first requirement semantic feature set. By transforming the creative intent explicitly expressed in natural language into executable control parameters, avoiding reliance solely on the text level, the controllability of the generation model or rendering engine is improved. This achieves the technical effect of providing a stable explicit control foundation for subsequent implicit semantic enhancement and batch generation of images and videos, thereby improving the accuracy of requirement parsing from the source.

[0046] Based on the semantic feature set of the advertising material generation requirements, batch iterative memory processing is performed on the image advertising material generation task to obtain batch generation results of image advertising materials. Based on the semantic feature set of advertising material generation requirements, batch iterative memory processing is performed on the video advertising material generation task to obtain the batch generation results of video advertising materials. Furthermore, based on the semantic feature set of the advertising material generation requirements, batch iterative memory processing is performed on the image advertising material generation task to obtain batch generation results of image advertising materials, including: A set of image advertising elements is constructed based on the set of semantic features of the generated advertising materials. A first image advertisement material is generated based on the set of image advertisement elements. First-order semantic parsing is performed on the first image advertisement material to obtain the semantic features of the first image advertisement material. The semantic features of the first image advertisement material are added to an initially empty vector to construct a first semantic memory unit. A second image advertisement material is generated based on the first semantic memory unit and the set of image advertisement elements; The first semantic memory unit is iteratively updated based on the second image advertising material to obtain the second semantic memory unit; The third image advertising material is generated based on the second semantic memory unit and the set of image advertising elements. This process is repeated until the preset batch generation quantity of the image advertising material is met, and the batch generation result of the image advertising material is obtained.

[0047] Furthermore, generating a second image advertisement material based on the first semantic memory unit and the set of image advertisement elements includes: Extract semantic features of the first image advertising material from the first semantic memory unit; Assign association weights to the semantic features of the first image advertisement material and the set of image advertisement elements to obtain a set of element association weights; The second image advertisement material is generated based on the set of associated weights of the elements and the set of image advertisement elements.

[0048] In one embodiment, the image ad element set refers to the set of basic constituent elements that can be used for image generation, decomposed from the semantic feature set of advertising material generation requirements. This typically includes main visual elements, such as product images and people images; auxiliary graphic elements, such as background textures and decorative lines; text elements, such as titles, subtitles, and calls to action; and style parameters, such as color ratios, white space ratios, and font types. First-order semantic parsing refers to extracting content-level features from the generated image, such as extracting its main color ratio, composition structure type, and text position distribution, and converting these into numerical semantic feature vectors.

[0049] Preferably, by loading the API interface specification of the image rendering engine, the visual rule features, layout matching features, and style control features in the semantic feature set of advertising material generation requirements are deconstructed and split according to the engine parameter format. The visual rule features are split into color parameter sets, texture parameter sets, and lighting parameter sets; the layout matching features are split into element positioning parameter sets, size ratio parameter sets, and alignment parameter sets; and the style control features are split into emotion target vectors, filter identifiers, and font style parameter sets, thereby obtaining the set of image advertising elements.

[0050] A pre-trained visual recognition model is used to perform first-order semantic parsing on the first image advertisement material to obtain features that reflect the material's characteristics, including visual symbol categories and locations. The visual recognition model is a target detection model based on the Faster R-CNN architecture, with its backbone network using a ResNet-50 pre-trained on the ImageNet dataset. The Faster R-CNN model consists of four core modules: the first module is a shared convolutional layer, which is the feature extraction part of the ResNet-50. The input image is pre-processed to a size of 224×224×3 before being fed into the ResNet-50 network. This network contains 49 convolutional layers and 1 fully connected layer, employing a Bottleneck residual block structure. It reduces computational complexity while maintaining feature representation capabilities through a combination of 1×1 convolutional dimensionality reduction, 3×3 convolutional feature extraction, and 1×1 convolutional dimensionality enhancement. The second module is a region extraction network. The network takes the feature map output from the shared convolutional layer as input and uses a 3×3 sliding window to densely sample the feature map. Each sliding window position has k pre-set anchor boxes of different scales and aspect ratios. The scale and aspect ratio of the anchor boxes are estimated and determined using the K-means clustering algorithm based on the size distribution of the labeled boxes in the training dataset. A typical configuration is a combination of 3 scales and 3 aspect ratios, resulting in 9 anchor boxes. The third module is the region of interest pooling layer. This layer maps the candidate regions output by the RPN onto the shared feature map and uses max pooling to uniformly sample candidate regions of different sizes into a fixed-size (e.g., 7×7) feature map to meet the fixed input dimension requirement of subsequent fully connected layers. The fourth module is the detection network. This network receives the fixed-size feature map output from the region of interest pooling layer and outputs two branches in parallel through several fully connected layers: the classification branch uses the Softmax function to calculate the probability distribution of candidate regions belonging to each category (including the background class); the regression branch further refines the bounding boxes of the candidate regions through regression, outputting the bounding box offset parameters corresponding to each category.

[0051] The ResNet-50 backbone network is pre-trained on the ImageNet large-scale image classification dataset. ImageNet contains approximately 1.28 million training images and 1000 class labels. The pre-training task is image classification, using the cross-entropy loss function and optimized by a stochastic gradient descent optimizer. After pre-training, the network has learned general visual feature representation capabilities such as edges, textures, and shapes. The pre-trained weights at this stage are publicly available and do not require repeated training in this step.

[0052] Based on a pre-trained ResNet-50, a complete Faster R-CNN detection network is constructed and fine-tuned end-to-end using an advertising creative domain dataset. The fine-tuning dataset consists of historical advertising creative images and their annotations. Annotations include visual symbol category labels, such as product cans, keyboards, screens, people, coffee cups, etc., along with their corresponding bounding box coordinates. The dataset size must meet the requirement of sample balance across categories, with a typical configuration of at least 500 annotated instances per category.

[0053] The fine-tuning process consists of four steps. Step 1: Training the RPN. Shared convolutional layers are initialized using ImageNet pre-trained weights. Newly added RPN layers are randomly initialized using a Gaussian distribution with a mean of 0 and a variance of 0.01. Using a single image as the basic unit for a mini-batch, 256 anchor boxes are randomly sampled from each image to form a mini-batch, maintaining a 1:1 ratio of positive to negative samples (positive samples are defined as anchor boxes with an IoU > 0.7 with the ground truth bounding boxes, negative samples are defined as anchor boxes with an IoU < 0.3, and anchor boxes with an IoU between 0.3 and 0.7 are not included in training). The SGD optimizer is used with an initial learning rate of 0.001, momentum of 0.9, and weight decay of 0.0005. After iteratively training approximately 60k mini-batches on the PASCAL VOC or a self-built advertising dataset, the learning rate is reduced to 0.0001 and iterated for another 20k times. Step 2: Based on the candidate regions generated by the RPN in Step 1, an independent Fast R-CNN detection network is trained. Similarly, ImageNet pre-trained weights are used for initialization, and newly added layers in the detection network are randomly initialized. During training, the RPN parameters are fixed, and only the detection network parameters are updated. Hyperparameter configuration is the same as in the first step. In the third step, the detection network weights trained in the second step are loaded, and the RPN is reinitialized using its shared convolutional layer parameters. The shared convolutional layer parameters are fixed and not updated; only the RPN-specific layers are fine-tuned. This step allows the RPN and the detection network to share the same feature extraction network. In the fourth step, the shared convolutional layer parameters are fixed, and only the fully connected layers of the detection network are fine-tuned. This completes the parameter sharing between the RPN and the detection network, forming a unified Faster R-CNN detection model, which is the aforementioned visual recognition model.

[0054] Next, the semantic features of the first image ad creative stored in the first semantic memory unit are read, and the values ​​of key dimensions are extracted, including visual symbol detection confidence, core symbol coordinate position, measured sentiment intensity, and measured layout distance. A cosine similarity-based association weight calculation module is loaded to calculate the correlation between the semantic features of the first image ad creative and each generated element in the image ad element set. For visual elements, the matching degree between the visual symbols actually appearing in historical creatives and the forced symbol set in the element set is calculated; symbols with successful matches receive high association weights. For layout elements, the offset between the measured coordinates of the core elements in historical creatives and the preset template coordinates in the element set is calculated; the smaller the offset, the higher the association weight. For style elements, the difference between the measured sentiment value in historical creatives and the target sentiment value in the element set is calculated; the larger the difference, the higher the adjustment weight for that element. The system outputs the association weight values ​​of each element, forming an element association weight set.

[0055] The set of element association weights is sent back to the parameter configuration module of the image rendering engine. The engine dynamically adjusts the set of image ad elements based on these weight values. High-weight visual symbols are marked as mandatory inheritance, meaning they must appear in the next image and their positional offset is constrained. High-weight layout parameters are set as baseline anchor points, and the positions of similar elements in subsequent materials must reference these anchor points. Style parameters corresponding to emotional differences are proportionally adjusted to make the emotional intensity of the new material approach the target trajectory. The adjusted parameter set is then merged with the image ad element set, driving the rendering engine to generate the second image ad material.

[0056] Furthermore, based on the same principle as obtaining the semantic features of the first image ad creative, first-order semantic parsing is performed on the second image ad creative to obtain its semantic features. The average of the semantic features of the first and second image ad creatives is calculated, and the result is used to replace the semantic features of the first image ad creative within the first semantic memory unit to obtain the second semantic memory unit. Similarly, a third image ad creative is generated based on the second semantic memory unit and the set of image ad elements. This process continues, performing creative generation and iterative updates of the semantic memory unit. If the number of iterations does not reach the preset batch generation quantity of creatives, the latest semantic memory unit and set of image ad elements are used as input to repeatedly execute the memory-guided generation - new creative parsing - memory unit iterative update loop until the preset batch generation quantity of creatives in the image ad creative generation task is met, thus obtaining the batch generation result of the image ad creatives.

[0057] Based on the same principle as obtaining batch generation results of image ad creatives, batch iterative analysis is performed during the batch generation of video ad creatives to ensure that the generated video ad creatives belong to the same category, thereby achieving the technical effect of improving the utilization rate of subsequent creatives.

[0058] The batch generation results of image ad creatives and video ad creatives are traversed and subjected to creative consistency checks. When the creative consistency check passes, the batch generation results of image ad creatives and video ad creatives are aggregated to obtain the batch generation results of the target ad creatives.

[0059] Furthermore, the batch generation results of the image ad creatives and the batch generation results of the video ad creatives are traversed and consistency checks are performed on the creatives, including: The batch generation results of image ad creatives and video ad creatives are traversed to extract multidimensional features, thereby obtaining multidimensional features of image ad creatives and multidimensional features of video ad creatives. The similarity between the multi-dimensional features of the image advertising material and the multi-dimensional features of the video advertising material is calculated to obtain the multi-dimensional feature similarity. When the multidimensional feature similarity is greater than or equal to a preset similarity threshold, the material consistency detection passes. When the similarity of the multidimensional features is less than the preset similarity threshold, the material consistency detection fails.

[0060] In one possible implementation, material consistency detection refers to the process of performing cross-modal semantic alignment verification on image and video ad creatives generated in batches within the same series. The detection object is not the internal quality of a single creative, but rather the external correspondence between the image and video sets. The main purpose is to ensure consistency in the content of the images and videos, avoiding inconsistencies between the generated ad images and videos.

[0061] Based on the aforementioned visual recognition model, feature extraction is performed on each video frame in the batch generation results of image ad creatives and video ad creatives, respectively. The extracted results are then averaged to obtain multi-dimensional features for both image and video ad creatives. Cosine similarity is calculated between the multi-dimensional feature vectors of the image and video ad creatives to obtain their multi-dimensional feature similarity. Cosine similarity measures the directional consistency of the angle between two vectors in high-dimensional space, with a value ranging from -1 to 1. A value closer to 1 indicates a more similar semantic orientation between the two vectors.

[0062] The calculated multidimensional feature similarity is compared with a preset similarity threshold set by those skilled in the art. The preset similarity threshold is a pre-configured hyperparameter, typically ranging from 0.75 to 0.90. When the similarity is greater than or equal to the preset similarity threshold, the material consistency detection is considered passed; when the similarity is less than the preset similarity threshold, the material consistency detection is considered failed, and the above steps need to be repeated, i.e., the material needs to be regenerated.

[0063] If the consistency check passes, the batch-generated image ad creatives and video ad creatives are packaged together by ad campaign identifier and platform dimension to generate a batch-generated target ad creative package that can be directly consumed by the ad delivery system. Traditional batch ad creative generation methods treat images and videos as two separate production lines, relying on manual quality control for visual consistency verification after output. This step, through automated multi-dimensional feature extraction and similarity calculation, transforms the subjective judgment of whether images and videos appear to belong to the same series into a quantifiable and reproducible numerical judgment problem, achieving an objective measurement of cross-modal consistency. This achieves the technical effect of improving the quality of batch-generated ad creatives.

[0064] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0065] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0066] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A system for batch generation of advertising creatives for images and videos, characterized in that, The system includes: The semantic enhancement processing module is used to obtain advertising material generation requirements, perform explicit and implicit semantic enhancement processing, and determine the set of semantic features of advertising material generation requirements; The image batch iterative memory processing module is used to generate a set of semantic features of demand based on the advertising material, perform batch iterative memory processing on the image advertising material generation task, and obtain the batch generation result of image advertising material; The video batch iterative memory processing module is used to generate a set of semantic features of demand based on the advertising material, perform batch iterative memory processing on the video advertising material generation task, and obtain the batch generation result of the video advertising material; The material consistency detection module is used to iterate through the batch generation results of image ad materials and video ad materials and perform material consistency detection on each. When the material consistency detection passes, the batch generation results of image ad materials and video ad materials are summarized to obtain the batch generation results of the target ad materials.

2. The system for batch generation of advertising materials for images and videos as described in claim 1, characterized in that, The semantic enhancement processing module also includes: The formatting processing unit is used to uniformly format the natural language creative briefing data, structured advertising configuration data, and platform specification data in the advertising material generation requirements, and construct the original data set of the requirements; An explicit semantic enhancement processing unit is used to perform explicit semantic enhancement processing on the original set of requirements to obtain a first set of requirements semantic features, wherein the first set of requirements semantic features includes visual rule features, layout matching features and style control features. An implicit semantic enhancement processing unit is used to perform implicit semantic enhancement processing on the original set of requirements data to construct a second set of requirements semantic features; The union unit is used to perform a union operation on the first set of demand semantic features and the second set of demand semantic features to determine the set of demand semantic features for generating advertising materials.

3. The system for batch generation of advertising materials for images and videos as described in claim 2, characterized in that, The implicit semantic enhancement processing unit also includes: The retrieval subunit is used to retrieve the historical advertising material repository using the original data set of the demand as an index, and obtain the historical advertising material set and the set of campaign performance data; The filtering subunit is used to filter the advertising effect data set according to preset advertising effect constraints to obtain an effective advertising effect data set. The mapping and extraction subunit is used to map and extract the historical advertising material set based on the effective delivery effect data set to obtain the effective historical advertising material set; The feature extraction subunit is used to extract implicit semantic features from the set of valid historical advertising materials to obtain an implicit semantic feature set. A semantic feature set is obtained by sub-units, which are used to perform implicit semantic enhancement processing on the implicit semantic feature set to obtain the second required semantic feature set.

4. The system for batch generation of advertising materials for images and videos as described in claim 3, characterized in that, The semantic feature set also includes the following sub-units: The feature mean processing micro-unit is used to traverse the implicit semantic feature set to perform feature mean processing and determine the initial implicit semantic enhancement center; The implicit semantic feature extraction micro-unit is used to extract implicit semantic features from the implicit semantic feature set whose Euclidean distance to the initial implicit semantic enhancement center satisfies the preset nearest neighbor bandwidth, and to construct the neighborhood of the initial implicit semantic enhancement center; The iterative implicit semantic enhancement center neighborhood construction micro-unit is used to take the implicit semantic features corresponding to the maximum Euclidean distance of the initial implicit semantic enhancement center neighborhood as the iterative implicit semantic enhancement center and construct the iterative implicit semantic enhancement center neighborhood. The stopping condition determination micro-unit is used to compare the size of the neighborhood features of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood to determine whether the preset enhancement processing stopping condition is met. If it is met, the initial implicit semantic enhancement center neighborhood is used as the second required semantic feature set. If the condition is not met, the implicit semantic features corresponding to the maximum Euclidean distance between the iterative implicit semantic enhancement center and the neighborhood of the iterative implicit semantic enhancement center are used to continue the enhancement process for the iterative implicit semantic enhancement center until the preset enhancement process stop condition is met.

5. The system for batch generation of advertising materials for images and videos as described in claim 4, characterized in that, When the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is greater than or equal to the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, the preset enhancement processing stop condition is met. When the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is less than or equal to a preset difference threshold, the preset enhancement processing stop condition is met. When the neighborhood feature quantity of the initial implicit semantic enhancement center neighborhood is less than the neighborhood feature quantity of the iterative implicit semantic enhancement center neighborhood, and the difference between the neighborhood feature quantities of the initial implicit semantic enhancement center neighborhood and the iterative implicit semantic enhancement center neighborhood is greater than a preset difference threshold, the preset enhancement processing stop condition is not met.

6. The system for batch generation of advertising materials for images and videos as described in claim 2, characterized in that, The explicit semantic enhancement processing unit also includes: The keyword extraction subunit is used to extract keywords from the original set of demand data to obtain a set of explicit demand keywords. The constraint-based filtering subunit is used to traverse the set of explicit requirement keywords from three dimensions: visual, layout, and style, and perform constraint-based filtering in the same dimension to obtain the set of explicit requirement keywords to be filtered. The structured encoding processing subunit is used to perform structured encoding processing on the set of explicit requirement keywords to obtain the first set of requirement semantic features.

7. The system for batch generation of advertising materials for images and videos as described in claim 1, characterized in that, The image batch iterative memory processing module also includes: The image ad element set construction unit is used to construct the image ad element set based on the ad material generating a set of semantic features of demand. A first-order semantic parsing unit is used to generate a first image advertisement material based on the set of image advertisement elements, perform first-order semantic parsing on the first image advertisement material to obtain the semantic features of the first image advertisement material, add the semantic features of the first image advertisement material into an initially empty vector, and construct a first semantic memory unit. The second image advertising material generation unit is used to generate a second image advertising material based on the first semantic memory unit and the image advertising element set; An iterative update unit is used to iteratively update the first semantic memory unit based on the second image advertising material to obtain the second semantic memory unit; The image ad creative batch generation result acquisition unit is used to generate a third image ad creative based on the second semantic memory unit and the image ad element set. This process continues until the preset batch generation quantity of the image ad creative is met, and the image ad creative batch generation result is obtained.

8. A batch generation system for advertising materials for images and videos as described in claim 7, characterized in that, The second image ad creative generation unit also includes: The first image advertising material semantic feature extraction subunit is used to extract the semantic features of the first image advertising material from the first semantic memory unit; The weight assignment subunit is used to assign associated weights to the semantic features of the first image advertising material and the set of image advertising elements to obtain a set of element associated weights. The material generation subunit is used to generate a second image advertisement material based on the element association weight set and the image advertisement element set.

9. A batch generation system for advertising materials for images and videos as described in claim 1, characterized in that, The material consistency detection module also includes: The multidimensional feature extraction unit is used to traverse the batch generation results of image ad creatives and video ad creatives to extract multidimensional features and obtain multidimensional features of image ad creatives and video ad creatives. The similarity calculation unit is used to calculate the similarity between the multi-dimensional features of the image advertising material and the multi-dimensional features of the video advertising material to obtain the multi-dimensional feature similarity. The detection pass unit is used to ensure that the material consistency detection passes when the multidimensional feature similarity is greater than or equal to a preset similarity threshold; The detection failure unit is used to prevent material consistency detection from failing when the multidimensional feature similarity is less than a preset similarity threshold.

10. A method for batch generation of advertising creatives for images and videos, characterized in that, The method is implemented by the batch generation system for advertising materials oriented towards images and videos as described in any one of claims 1-9, and the method includes: Obtain the advertising creative generation requirements, perform explicit and implicit semantic enhancement processing, and determine the semantic feature set of the advertising creative generation requirements; Based on the semantic feature set of the advertising material generation requirements, batch iterative memory processing is performed on the image advertising material generation task to obtain batch generation results of image advertising materials. Based on the semantic feature set of advertising material generation requirements, batch iterative memory processing is performed on the video advertising material generation task to obtain the batch generation results of video advertising materials. The batch generation results of image ad creatives and video ad creatives are traversed and subjected to creative consistency checks. When the creative consistency check passes, the batch generation results of image ad creatives and video ad creatives are aggregated to obtain the batch generation results of the target ad creatives.