Commodity marketing picture and video generation method and device, equipment and medium

By introducing structured product semantic tags and multimodal models, and integrating the product marketing image and video generation process, the problem of visual and style inconsistencies caused by separate processing is solved, and efficient and automated product marketing material generation is achieved.

CN121414908APending Publication Date: 2026-01-27BEIJING WEZONET NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511530407.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

In existing technologies, the separation of product marketing image and video generation processes leads to poor consistency in visual context, style, and brand tone. Furthermore, the process is highly complex, cumbersome, and requires manual intervention at multiple stages and with multiple tools.

Method used

By introducing structured product semantic tags as a unified context throughout the process, and using a multimodal model to understand product content, generate creative content, and collaboratively process chronological content, an end-to-end process is integrated to generate product marketing materials with consistent visual content and style.

Benefits of technology

It improves the efficiency of creating product marketing materials, reduces the professional skills required of users, and ensures consistency in visual context, style, and brand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414908A_ABST
    Figure CN121414908A_ABST
Patent Text Reader

Abstract

The invention discloses a commodity marketing picture and video generation method and device, equipment and a medium. The method comprises the following steps: receiving an input commodity image and description information provided by a user; carrying out preprocessing and main body segmentation on the commodity image to obtain a commodity main body image; based on the commodity main body image and the user description information, commodity content understanding is carried out through a multi-modal model, a structured commodity semantic tag is generated, and the commodity semantic tag comprises the category, the visual feature and the key attribute of the commodity; and on the basis of the commodity semantic label, creative content generation and automatic typesetting processing are carried out through a multi-modal model, and a commodity marketing image containing a background, an advertisement copywriting and decoration elements is generated. According to the method, the tedious process needing multiple stages, multiple tools and repeated manual intervention is integrated into a highly-automatic end-to-end process, the manufacturing efficiency of commodity marketing materials is improved, meanwhile, the requirement for professional skills of a user is lowered, and the consistency of the visual context, style and brand is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to methods, apparatus, equipment and media for generating product marketing images and videos. Background Technology

[0002] With the rapid development of deep learning and generative artificial intelligence technologies, generating high-quality product marketing images and videos using multimodal models has become an important means in e-commerce and advertising marketing.

[0003] Existing technologies typically employ a phased, multi-model collaborative workflow, where image generation and video generation are independent steps. For example, the image generation stage might use a text-to-image model, while the video generation stage is based on a different model and prompts. This separate processing results in a lack of unified contextual association between images and videos, making it difficult to maintain consistency in visual details, style, and brand tone. Furthermore, users need to repeatedly design and optimize prompts at different stages, making the process cumbersome and lacking in controllability, increasing operational complexity and time costs. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, device, and medium for generating product marketing images and videos, aiming to solve the problem that the output content has significant differences in visual context, style, and brand consistency due to the separate processing of image and video generation.

[0005] In a first aspect, embodiments of the present invention provide a method for generating product marketing images and videos, comprising:

[0006] Receive the input product image and the description information provided by the user; The product image is preprocessed and segmented to obtain the main product image; Based on the product image and user description information, a multimodal model is used to understand the product content and generate structured product semantic tags. The product semantic tags include the product category, visual features and key attributes. Based on the product semantic tags, creative content is generated and automatically formatted using a multimodal model to generate product marketing images that include backgrounds, advertising copy, and decorative elements. Based on the product semantic tags, product marketing images, and descriptive information, a multimodal model is used to perform time-series content collaborative generation processing to generate a product marketing video that is consistent with the product marketing images in terms of visual content and style.

[0007] In a second aspect, embodiments of the present invention provide a product marketing image and video generation device, comprising: The receiving unit is used to receive the input product image and the description information provided by the user; The image processing unit is used to preprocess and segment the product image to obtain a product subject image; The tag generation unit is used to generate structured semantic tags for the product based on the product image and user description information, through multimodal model to understand the product content. The semantic tags include the product category, visual features and key attributes. The image generation unit is used to generate creative content and automatically typeset based on the product semantic tags through a multimodal model, and generate a product marketing image that includes background, advertising copy and decorative elements; The video generation unit is used to generate a product marketing video that is consistent with the product marketing image in terms of visual content and style by performing time-series content collaborative generation processing based on the product semantic tags, product marketing images and descriptive information through a multimodal model.

[0008] Thirdly, embodiments of the present invention provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the product marketing image and video generation method described in the first aspect.

[0009] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the product marketing image and video generation method described in the first aspect.

[0010] The beneficial effects of this invention are as follows: By introducing structured product semantic tags as a unified context throughout the process, the problem of contextual fragmentation and consistency caused by the phased generation of images and videos is fundamentally solved. This device integrates the originally cumbersome process requiring multiple stages, multiple tools, and repeated manual intervention into a highly automated end-to-end process, significantly improving the efficiency of product marketing material production while reducing the professional skills required of users and ensuring consistency in visual context, style, and brand. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1This is a flowchart illustrating the method for generating product marketing images and videos provided in an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of a sub-process of step S102 provided in an embodiment of the present invention.

[0014] Figure 3 This is a schematic diagram of a sub-process of step S103 provided in an embodiment of the present invention.

[0015] Figure 4 This is a schematic diagram of a sub-process of step S104 provided in an embodiment of the present invention.

[0016] Figure 5 This is a schematic diagram of a sub-process of step S105 provided in an embodiment of the present invention.

[0017] Figure 6 This is a schematic block diagram of a product marketing image and video generation device provided in an embodiment of the present invention.

[0018] Figure 7 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0021] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0022] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0023] Please see Figure 1 , Figure 1 A flowchart illustrating the method for generating product marketing images and videos provided in an embodiment of the present invention; like Figure 1 As shown, the method includes steps S101-S105.

[0024] S101, Receive the input product image and user-provided description information; In step S101, the description information refers to the information on the product image that covers the key features of the product, such as the product's category, color, size, material, function, and other details.

[0025] S102. Preprocess and segment the product image to obtain the product subject image; Step S102 aims to extract a clean image of the main product through preprocessing and subject segmentation techniques, thereby eliminating the interference of complex backgrounds on subsequent steps.

[0026] S103. Based on the product's main image and user description information, a multimodal model is used to understand the product content and generate structured product semantic tags. The product semantic tags include the product's category, visual features, and key attributes. Step S103 aims to use a multimodal model to perform in-depth content understanding of the product subject and generate a structured product semantic tag. This tag systematically summarizes the product category, visual features and key attributes, providing a unified semantic context for the entire process.

[0027] S104. Based on the semantic tags of the product, creative content is generated and automatically formatted using a multimodal model to generate a product marketing image that includes background, advertising copy and decorative elements; Step S104 aims to trigger the creative content generation and automatic layout process based on the semantic tags of the product, and produce a product marketing image that integrates customized background, advertising copy and decorative elements.

[0028] S105. Based on product semantic tags, product marketing images and descriptive information, a multimodal model is used to perform time-series content collaborative generation processing to generate a product marketing video that is consistent with the product marketing images in terms of visual content and style.

[0029] Step S105 aims to utilize the aforementioned generated product semantic tags, product marketing images, and descriptive information to ensure that the final output product marketing video maintains a high degree of consistency with the previous images in terms of subject, style, and scene through time-series content collaborative generation technology.

[0030] In this embodiment, steps S101-S105 fundamentally solve the problem of contextual fragmentation and consistency caused by the phased generation of images and videos by introducing structured product semantic tags as a unified context throughout the process. This embodiment integrates the originally cumbersome process that required multiple stages, multiple tools, and repeated manual intervention into a highly automated end-to-end process, significantly improving the efficiency of product marketing material production, while reducing the professional skills required of users and ensuring consistency in visual context, style, and brand.

[0031] In one embodiment, such as Figure 2 As shown, step S102 includes: S201. Perform resolution enhancement and noise removal processing on the product image; S202. Input the preprocessed product image into the target segmentation model based on deep learning for segmentation processing, and output the pixel-level segmentation mask of the main body of the product in the product image. S203. Based on pixel-level segmentation mask, extract the main body of the product from the product image to obtain the main body image of the product.

[0032] In this embodiment, firstly, by performing resolution enhancement and noise removal on the product image, the image quality can be improved, providing a foundation for subsequent accurate segmentation. Then, a deep learning-based target segmentation model, such as the Mask R-CNN instance segmentation model, is employed. The Mask R-CNN instance segmentation model uses its region proposal network to first locate candidate regions in the product image that may contain the product. Then, each candidate region is classified and its precise pixel-level segmentation mask is predicted. The pixel-level segmentation mask is a binary matrix of the same size as the original product image, where each pixel in the main product region is precisely labeled as foreground (e.g., a value of 1), and the remaining background parts are labeled as background (e.g., a value of 0), thus achieving a pixel-level definition of the product outline. Finally, based on the pixel-level segmentation mask, the main product is extracted from the product image to obtain the main product image.

[0033] For example, taking the segmentation of a pair of sneakers as an example, Mask R-CNN can accurately delineate the boundaries of complex structures such as shoelaces, uppers, and soles. Finally, based on this pixel-level segmentation mask, image processing techniques are used to losslessly separate the main product from the product image, generating a product image with a transparent background (alpha channel). This provides great flexibility for subsequent creative compositing by freely placing it in various AI-generated backgrounds.

[0034] In one embodiment, step S202 is followed by: performing morphological post-processing on the pixel-level segmentation mask to optimize the smoothness and integrity of the product body edge.

[0035] In this embodiment, morphological post-processing effectively repairs subtle imperfections that may arise from the segmentation model, generating a product subject mask with smoother edges and a more complete internal structure. This improves the quality of the final extracted product subject image, making its edge transitions more natural. It is particularly beneficial for achieving seamless integration of the subject and the synthetic background when generating high-quality marketing images and videos, thus enhancing the image's texture.

[0036] In one embodiment, such as Figure 3 As shown, step S103 includes: S301. Visual feature extraction and semantic analysis of the product image are performed using a multimodal model to output a set of keywords to describe the product content. S302. Merge and structure the keyword set according to the description information to generate a semantic tag for the product containing the product category, visual features and key attributes; wherein, the key attributes include functional attributes, physical attributes and scene information.

[0037] In this embodiment, the multimodal model in step S301 can be a visual language model capable of simultaneously processing and understanding image and text information, such as the BLIP-2 model developed by Google. The BLIP-2 model uses a visual Transformer (ViT) as an image encoder to encode the main image of the product into a series of visual feature vectors; at the same time, it uses a pre-trained language model (such as Flan-T5) as a text encoder and decoder.

[0038] In step S301, the visual encoder of the BLIP-2 model encodes the main image of the product into a patch embedding sequence, and extracts visual features through a multi-layer self-attention mechanism. Its Q-Former module performs cross-attention calculation between the learnable query vector and the visual features, outputting a visual semantic embedding. The text decoder of the BLIP-2 model then generates a set of descriptive keywords based on this visual semantic embedding.

[0039] For step S302, the text decoder of the BLIP-2 model performs semantic analysis on the user description information to obtain text features. The text features and visual semantics are embedded in the feature space for alignment and fusion. Then, the attention mechanism is used to achieve keyword deduplication and summarization, and finally, structured product semantic tags are output.

[0040] Based on this, steps S301-S302 of this embodiment, by employing a visual language model such as BLIP-2, can achieve high-precision visual feature extraction and semantic understanding of product images, transforming unstructured visual and textual information into structured data that machines can accurately understand and consistently access. This transformation enables the core semantics of the product to be stably and unambiguously transmitted at different generation stages (such as text generation and video generation), which is a key technology for solving the image-visual consistency problem, and also provides a rich and standardized input source for subsequent automated creative generation.

[0041] In one embodiment, such as Figure 4 As shown, step S104 includes: S401. Input the category, visual features and key attributes in the structured product semantic tags into a large language model for copy matching processing, and output advertising and marketing copy that matches the tags. S402. Input the main image of the product, the semantic tags of the product, and the style instructions that the user can select into the Wenshengtu multimodal diffusion model, perform an end-to-end generation process of background fusion generation, product fusion and stylized rendering, and output the initial marketing image. S403. Combine the initial marketing image with the advertising copy, process the image and text layout through an automated layout engine, and add system-preset or dynamically generated decorative elements. Finally, output a product marketing image that includes background, advertising copy and decorative elements.

[0042] For step S401, a large language model (such as GPT-4) can be used to generate precise advertising and marketing copy based on structured product semantic tags.

[0043] In step S402, the product image, product semantic tags, and user-selectable style instructions are input into the text-based image multimodal diffusion model (e.g., the Stable Diffusion XL model released by Stability AI), and its ControlNet or T2I-Adapter control network plugins are used to maintain the shape of the product. Next, the product semantic tags are used as text prompts, and the product image is transformed into control conditions by the encoder of the ControlNet plugin (e.g., Canny edge detection or depth map encoder). This guides the text-based image multimodal diffusion model to maintain the original shape and details of the product while generating the background based on the text semantics, achieving a natural blending of the product and the generated background. The user-selectable style instructions (e.g., selecting a watercolor style) influence the stylized rendering of the entire generation process through a cross-attention layer. Thus, the initial marketing image can be output through a single forward propagation denoising process.

[0044] For step S403, the initial marketing image and the generated advertising copy are synthesized by the layout engine, and the layout of the image and text, font matching and decorative elements are automatically completed based on image content analysis and design rules. Finally, a product marketing image containing background, advertising copy and decorative elements is output.

[0045] In this embodiment, the collaborative work of a large-scale language model and a text-to-image multimodal diffusion model achieves a high-quality conversion from semantic understanding to visual generation, ensuring a high degree of consistency between creative content and product characteristics. In particular, the application of ControlNet technology precisely controls the morphological features of the product subject while maintaining creative freedom, avoiding deformation problems common in traditional methods. Combined with subsequent automated typesetting, a complete automated process from product understanding to finished product output is formed, effectively improving the efficiency and quality of marketing image production.

[0046] In one embodiment, step S403 includes: The initial marketing image is synthesized with the advertising copy, and the main product area and visual focus area in the initial marketing image are identified. Based on the identified main product area and visual focus area, the placement of advertising and marketing copy is calculated, with the principle of ensuring that the copy content is not obscured by the main product area and visual focus area. Based on the style information in the product's semantic tags, the system automatically matches the corresponding font, color, and size to the advertising copy, adds system-preset or dynamically generated decorative elements, and outputs a product marketing image that includes background, advertising copy, and decorative elements. Based on the different advertising platform deployment requirements, product marketing images are rendered into versions with various preset sizes and proportions.

[0047] In this embodiment, firstly, the initial marketing image and advertising copy are synthesized. Then, the GrabCut algorithm can be used for product subject region segmentation. Specifically, the GrabCut algorithm can accurately separate the product subject region and the background region using foreground / background modeling. Simultaneously, a deep learning-based saliency detection model (such as U...) can be used. 2 -Net calculates the visual saliency map of the image and identifies the visual focus areas that attract the most human eye. This yields the main product area and the visual focus areas.

[0048] Next, the main product area and the visual focus area are defined as obstacles with strong repulsive forces, and the advertising copy is modeled as a movable unit affected by these repulsive forces. The optimal placement of the advertising copy in the image is calculated by simulating the repulsive force and boundary constraints to ensure that the copy content neither obscures the main product nor covers the visual focus area, thereby maximizing the visual appeal and information delivery efficiency of the advertisement.

[0049] Next, the style information in the semantic tags of the product is analyzed, and a font type that coordinates with it is selected from the preset font library; at the same time, the k-means color clustering algorithm can be used to extract the dominant color system from the main image of the product, and the text color that forms the best visual effect is calculated based on the color contrast theory; the addition of decorative elements can be achieved through template matching and adaptive scaling technology to ensure that elements such as brand logos and borders are consistent with the overall design style.

[0050] Finally, a bilinear interpolation algorithm can be used; based on the different advertising platform specifications (such as Instagram's 1:1 square, Facebook's 16:9 banner, etc.), by maintaining the aspect ratio through adaptive cropping and intelligent edge-filling strategies, multiple sizes of marketing images can be generated to ensure the best visual effect in different display scenarios.

[0051] Based on this, the automated typesetting solution in this embodiment achieves full automation of the typesetting design process. It can intelligently avoid visual conflicts, accurately match design styles, and generate multi-format outputs in batches, improving the production efficiency and professional consistency of marketing images, while ensuring visual optimization for cross-platform display.

[0052] In one embodiment, such as Figure 5 As shown, step S105 includes: S501. Align and merge product semantic tags, product marketing images and descriptive information to construct a unified text-image joint prompt sequence that includes static visuals and dynamic intents; S502. Input the joint cue sequence into the Wensheng video multimodal diffusion model, perform temporal coherence generation processing with the product marketing image as the visual reference and the joint cue sequence as the cross-modal constraint, and finally output a dynamic marketing video that is consistent with the product marketing image in terms of product subject, visual style and scene context.

[0053] For step S501, an image encoder can be used to convert the product marketing image into a visual embedding vector; at the same time, the product semantic labels and description information can be converted into text embedding vectors through a text encoder; then, a feature alignment network is used to concatenate the visual embedding vectors and text embedding vectors in the latent space and perform layer normalization to construct a unified text-image joint cue sequence, which contains both accurate visual references and rich semantic guidance.

[0054] For step S502, the multimodal diffusion model can adopt the Make-A-Video model. In each denoising step, this model can simultaneously handle intra-frame spatial relationships and inter-frame temporal coherence through its spatiotemporal attention mechanism. Specifically, the model's spatial self-attention layer ensures that visual elements within a single frame conform to the semantic requirements of the joint cue sequence, while the temporal attention layer maintains motion smoothness and content consistency between frames. Therefore, this model uses the product marketing image as the first frame visual reference and generates subsequent video frames through a progressive denoising process. During the generation of each frame, the cross-modal constraints of the joint cue sequence are followed, ensuring that the generated dynamic marketing video maintains a high degree of consistency with the product marketing image in terms of product details, visual style elements, and scene context.

[0055] In this embodiment, based on the solutions in steps S501-S502, the Make-A-Video model is used to convert static product marketing images into dynamic marketing videos. While maintaining the creativity of the generated content, it ensures the visual and semantic consistency between the video sequence of the dynamic marketing video and the product marketing images. This ability to collaboratively generate temporal content effectively solves the common image-visual inconsistency problem in traditional methods, providing product marketing with truly dynamic materials that achieve image-visual consistency.

[0056] This invention also provides a product marketing image and video generation apparatus, which is used to execute any of the aforementioned product marketing image and video generation methods. Specifically, please refer to... Figure 6 , Figure 6 This is a schematic block diagram of the product marketing image and video generation device provided in the embodiments of the present invention.

[0057] like Figure 6 As shown, the product marketing image and video generation device 600 includes: a receiving unit 601, an image processing unit 602, a tag generation unit 603, an image generation unit 604, and a video generation unit 605.

[0058] The receiving unit 601 is used to receive the input product image and the description information provided by the user; Image processing unit 602 is used to preprocess and segment the product image to obtain the product main image; The tag generation unit 603 is used to generate structured semantic tags for products based on the product's main image and user description information through a multimodal model to understand the product content. The semantic tags for products include the product's category, visual features, and key attributes. The image generation unit 604 is used to generate creative content and automatically typeset based on product semantic tags through a multimodal model, generating product marketing images that include backgrounds, advertising copy, and decorative elements; The video generation unit 605 is used to generate a product marketing video that is consistent with the product marketing image in terms of visual content and style by using a multimodal model to collaboratively generate temporal content based on product semantic tags, product marketing images and descriptive information.

[0059] This device fundamentally solves the problem of contextual fragmentation and consistency caused by the phased generation of images and videos by introducing structured product semantic tags as a unified context throughout the process. It integrates the previously cumbersome process, which required multiple stages, multiple tools, and repeated manual intervention, into a highly automated end-to-end workflow. This significantly improves the efficiency of producing product marketing materials while reducing the professional skills required of users, ensuring the consistency and professionalism of the brand's visual output.

[0060] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0061] The aforementioned product marketing image and video generation device can be implemented in the form of a computer program, which can, for example, Figure 7 It runs on the computer device shown.

[0062] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device 700 is a server, which can be a standalone server or a server cluster composed of multiple servers.

[0063] See Figure 7 The computer device 700 includes a processor 702, a memory, and a network interface 705 connected via a system bus 701. The memory may include a non-volatile storage medium 703 and internal memory 704.

[0064] The non-volatile storage medium 703 may store an operating system 7031 and a computer program 7032. When the computer program 7032 is executed, it causes the processor 702 to execute a method for generating product marketing images and videos.

[0065] The processor 702 provides computing and control capabilities to support the operation of the entire computer device 700.

[0066] The internal memory 704 provides an environment for the operation of the computer program 7032 in the non-volatile storage medium 703. When the computer program 7032 is executed by the processor 702, the processor 702 can execute a method for generating product marketing images and videos.

[0067] This network interface 705 is used for network communication, such as providing data transmission. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device 700 to which the present invention is applied. The specific computer device 700 may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0068] Those skilled in the art will understand that Figure 7 The embodiments of the computer device shown do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those shown. Figure 7 The embodiments shown are consistent and will not be described again here.

[0069] It should be understood that, in this embodiment of the invention, the processor 702 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0070] In another embodiment of the invention, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the product marketing image and video generation method of the embodiments of the present invention.

[0071] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code.

[0072] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0073] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating product marketing images and videos, characterized in that, include: Receive the input product image and the description information provided by the user; The product image is preprocessed and segmented to obtain the main product image; Based on the product image and user description information, a multimodal model is used to understand the product content and generate structured product semantic tags. The product semantic tags include the product category, visual features and key attributes. Based on the product semantic tags, creative content is generated and automatically formatted using a multimodal model to generate product marketing images that include background, advertising copy, and decorative elements. Based on the product semantic tags, product marketing images, and descriptive information, a multimodal model is used to perform time-series content collaborative generation processing to generate a product marketing video that is consistent with the product marketing images in terms of visual content and style.

2. The method for generating product marketing images and videos according to claim 1, characterized in that, The step of preprocessing and subject segmentation of the product image to obtain the product subject image includes: The product image is subjected to resolution enhancement and noise removal processing; The preprocessed product image is input into a deep learning-based target segmentation model for segmentation, and the pixel-level segmentation mask of the main product in the product image is output. Based on the pixel-level segmentation mask, the main body of the product is extracted from the product image to obtain the main body image of the product.

3. The method for generating product marketing images and videos according to claim 2, characterized in that, The step of inputting the preprocessed product image into a deep learning-based target segmentation model for segmentation processing, and outputting a pixel-level segmentation mask of the main product in the product image, includes: Morphological post-processing is performed on the pixel-level segmentation mask to optimize the smoothness and integrity of the product body edge.

4. The method for generating product marketing images and videos according to claim 1, characterized in that, The step of generating structured product semantic tags based on the product image and user description information, using a multimodal model to understand product content, includes: The product image is subjected to visual feature extraction and semantic analysis using a multimodal model, and a set of keywords is output to describe the content of the product. The keyword set is merged and structured based on the description information to generate a product semantic tag containing the product category, visual features, and key attributes; wherein, the key attributes include functional attributes, physical attributes, and scene information.

5. The method for generating product marketing images and videos according to claim 1, characterized in that, The process of generating product marketing images based on the product semantic tags, using a multimodal model for creative content generation and automatic layout processing, and including background, advertising copy, and decorative elements, includes: The category, visual features, and key attributes in the structured product semantic tags are input into a large language model for copy matching processing, and the output is an advertising and marketing copy that matches the tags. The product image, product semantic tags, and user-selectable style instructions are input into the Wenshengtu multimodal diffusion model, and an end-to-end generation process of background fusion generation, product fusion, and stylized rendering is executed to output the initial marketing image. The initial marketing image is combined with the advertising copy, and the layout is processed by an automated typesetting engine. Pre-set or dynamically generated decorative elements are added, and the final output is a product marketing image that includes a background, advertising copy, and decorative elements.

6. The method for generating product marketing images and videos according to claim 4, characterized in that, The process of combining the initial marketing image with the advertising copy, performing image and text layout processing through an automated typesetting engine, and adding system-preset or dynamically generated decorative elements, ultimately outputting a product marketing image containing a background, advertising copy, and decorative elements, includes: The initial marketing image is combined with the advertising copy, and the main product area and visual focus area in the initial marketing image are identified. Based on the identified main product area and visual focus area, the placement position of the advertising copy is calculated, wherein the placement position is based on the principle of ensuring that the copy content is not obscured by the main product area and visual focus area. Based on the style information in the product semantic tags, the system automatically matches the corresponding font, color and size to the advertising copy, and adds system-preset or dynamically generated decorative elements to output a product marketing image that includes background, advertising copy and decorative elements. Based on the different advertising platform deployment requirements, the product marketing images are rendered into versions with various preset sizes and proportions.

7. The method for generating product marketing images and videos according to claim 1, characterized in that, The step of generating a product marketing video that is visually and stylistically consistent with the product marketing image, based on the product semantic tags, product marketing images, and descriptive information, through a multimodal model for time-series content collaborative generation, includes: The product semantic tags, product marketing images, and description information are aligned and fused to construct a unified text-image joint prompt sequence that includes static visuals and dynamic intents; The joint cue sequence is input into the Wensheng video multimodal diffusion model, and temporal coherence generation processing is performed with the product marketing image as the visual reference and the joint cue sequence as the cross-modal constraint condition. Finally, a dynamic marketing video that is consistent with the product marketing image in terms of product subject, visual style and scene context is output.

8. A device for generating product marketing images and videos, characterized in that, include: The receiving unit is used to receive the input product image and the description information provided by the user; The image processing unit is used to preprocess and segment the product image to obtain a product subject image; The tag generation unit is used to generate structured semantic tags for the product based on the product image and user description information, through multimodal model to understand the product content. The semantic tags include the product category, visual features and key attributes. The image generation unit is used to generate creative content and automatically typeset based on the product semantic tags through a multimodal model, and generate a product marketing image that includes background, advertising copy and decorative elements; The video generation unit is used to generate a product marketing video that is consistent with the product marketing image in terms of visual content and style by performing time-series content collaborative generation processing based on the product semantic tags, product marketing images and descriptive information through a multimodal model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating product marketing images and videos as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the product marketing image and video generation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • An advertisement material layout generation method and device based on artificial intelligence and a medium

    CN122435076A