Artificial intelligence-based product marketing short video generation method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
然而,在面向高转化要求的产品营销场景时,这些通用或特定领域的AIGC视频生成方案存在系统性的技术局限
通过建立人物-产品-营销的三维协同机制,执行基于产品属性标签、营销场景标签及目标受众画像的加权融合计算,实现了人物风格类型与动作姿势序列的精准匹配,提升了视频内容的营销适配性与带货说服力;通过构建产品核心卖点视觉区域的智能识别机制与场景化展示方案,将识别出的核心卖点区域与预置的场景素材库进行融合,生成产品与使用场景自然结合的展示方案,并依据营销逻辑自动规划展示顺序与时长分配,实现了产品展示精准度与吸引力的实质性提升;通过建立定制化音频与视频镜头序列的协同匹配机制,解决了现有方案中音视频内容相互独立、营销协同性差的问题;通过引入预设的时序编排模板及企业级定制元素的模块化整合机制,实现了从产品引入到购买引导的完整营销链路闭环,并满足了品牌差异化表达的企业级深度定制需求。
Smart Images

Figure CN122554699A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a method and device for generating short product marketing videos based on artificial intelligence. Background Technology
[0002] With the booming development of e-commerce live streaming and content marketing, short videos have become a key vehicle for product traffic generation and conversion. Advances in AI-generated content technology have provided fundamental support for the automated creation of short videos, with existing solutions generally following a basic paradigm of material input – template matching – automatic synthesis. Represented by patent document CN121691849A, this solution obtains user-inputted product information and style parameters, retrieves matching target template scripts from a dynamic template knowledge base, and then inputs the script content and product information into a large language model to generate the final video script. This template-matching-based generation method automates the script level and can update template weights based on campaign data feedback, thus improving the conversion rate of script generation to some extent. However, when facing product marketing scenarios with high conversion requirements, these general or domain-specific AIGC video generation solutions have systemic technical limitations.
[0003] First, at the content generation level, this solution only matches template scripts using discrete tags such as style parameters, lacking a multi-dimensional, refined matching mechanism for product attributes, marketing scenarios, and target audiences. This results in poor compatibility between the style and actions of the characters in the generated videos and the product characteristics and usage scenarios. The core selling points of the product cannot be effectively highlighted visually, and there is a disconnect between the audio explanation and the video presentation, making it difficult to construct a persuasive marketing narrative. Second, at the product presentation level, the existing solution only focuses on generating script content. The generation logic is not embedded in the complete product marketing chain, and it cannot automatically achieve a closed loop from product introduction to purchase guidance, failing to solve the problem of the product's core selling points not being accurately highlighted visually. Summary of the Invention
[0004] This application provides a method and device for generating short product marketing videos based on artificial intelligence to solve the above-mentioned technical problems.
[0005] On the one hand, embodiments of this application provide a method for generating short product marketing videos based on artificial intelligence, including: The system receives and standardizes user-inputted customized demand data. Based on the processed customized demand data, it performs a three-dimensional weighted fusion calculation of product attribute tag matching basic association weight, marketing scenario tag matching scenario adaptation coefficient, and target audience profile to calculate audience preference matching degree. It then performs person-product marketing adaptation processing to generate an adapted person display scheme. The customized demand data includes person feature parameters, product information data, and marketing configuration parameters. Based on the product information data and the marketing configuration parameters, the core selling points visual areas of the product are identified and extracted. Based on the identification results and the pre-set scene material library, a product display scheme that integrates the product and the usage scenario is generated. Based on the character display scheme, the product display scheme, the product information data, and the marketing configuration parameters, a customized audio is generated that matches the rhythm of the video content and the marketing logic, and a collaborative matching relationship is established between the customized audio and the video shot sequence. Based on a preset time-series arrangement template, the character display scheme, the product display scheme, and the customized audio are automatically arranged and integrated to generate a preliminary synthesized video.
[0006] In one implementation of this application, based on the processed customized demand data, a three-dimensional weighted fusion calculation is performed using product attribute tag matching basic association weights, marketing scenario tag matching scenario adaptation coefficients, and target audience profiles to calculate audience preference matching degree. This process is then used to perform person-product marketing adaptation processing and generate an adapted person display scheme, specifically including: Based on product attribute tags from product information data, the basic association weights for matching are obtained from a preset adaptation rule base. Based on the marketing scenario tags in the marketing configuration parameters, the matching scenario adaptation coefficient is obtained from the adaptation rule base, and based on the target audience profile in the marketing configuration parameters, the audience preference matching degree is calculated from the adaptation rule base. The basic association weight, the scene adaptation coefficient, and the audience preference matching degree are weighted and fused together to generate the final adaptation weight; Based on the ranking results of the final adaptation weights, the optimal character style type and action posture sequence are selected from the adaptation rule base and combined to generate an adapted character display scheme.
[0007] In one implementation of this application, based on the product information data and the marketing configuration parameters, the core selling point visual area of the product is identified and extracted, specifically including: The product images in the product information data are preprocessed to extract multi-scale basic image features, and the product attribute tags in the product information data are semantically encoded to generate product attribute semantic vectors. Based on the product attribute semantic vector, semantic attention weights are calculated on the basic features of the image to generate an attention weight feature map that focuses on the semantically related areas of the product's selling points. Based on the attention weight feature map, the target detection model is driven to identify product images and output core selling point visual region data; the core selling point visual region data includes the coordinates, category and confidence level of the core selling point region.
[0008] In one implementation of this application, based on the character display scheme, the product display scheme, the product information data, and the marketing configuration parameters, customized audio is generated to match the rhythm of the video content and the marketing logic, specifically including: Based on the product selling point text in the product information data, the marketing scenario rules in the marketing configuration parameters, and the target audience preferences, a segmented script is generated using a pre-trained language model, and a segmented speech script with a marketing logic structure is output. Based on the segmented speech script and the character style type in the character display scheme, the corresponding emotional speech synthesis model is called to generate segmented speech audio with timestamps. Based on the marketing scenario rules in the marketing configuration parameters and the overall rhythm of the segmented voice script, the corresponding background music segments are matched and extracted from the background music library; The segmented audio and background music clips are then mixed and balanced in volume to generate the final customized audio.
[0009] In one implementation of this application, establishing the collaborative matching relationship between the customized audio and video shot sequences specifically includes: The timestamp markers in the customized audio are parsed to obtain audio content segmentation nodes, and the video shot sequence is planned according to the time sequence arrangement template to obtain video shot segmentation nodes; Based on the preset mapping rules, the audio content segment nodes are aligned and matched with the video shot segment nodes to generate an initial audio-visual synchronization alignment scheme. The duration of each segment of the audio content segment node in the initial audio-visual synchronization alignment scheme is compared with the duration of the corresponding segment of the video shot segment node in real time, and the duration deviation is calculated. When the duration deviation exceeds a preset threshold, the playback rate of the video shot sequence or the speech rate of the customized audio is dynamically adjusted to generate an adjusted audio-visual synchronization alignment scheme to achieve audio-visual synchronization.
[0010] In one implementation of this application, the automated arrangement and integration of the character display scheme, the product display scheme, and the customized audio based on a preset temporal arrangement template to generate a preliminary synthesized video specifically includes: The character display scheme, the product display scheme, and the customized audio are mapped and associated with each stage in the preset time sequence arrangement template according to the content attributes to generate the original content mapping relationship. The time sequence arrangement template is a structured configuration file that defines the stage order, duration range, and content type constraints. It includes a first time sequence segment, a second time sequence segment, a third time sequence segment, and a fourth time sequence segment. Each segment corresponds to different content type combinations and different time windows with durations greater than a threshold. Based on the target audience and promotional information in the marketing configuration parameters, the original content mapping relationship data is personalized with content filling and duration allocation to generate the arranged video content structure; The customized audio is segmented and aligned according to the video content structure, and the character display scheme and the product display scheme are rendered into video streams. They are then integrated according to the video content structure to generate a preliminary composite video.
[0011] In one implementation of this application, receiving and standardizing user-input customized request data specifically includes: The system collects user-defined character image, gender, age, style, and action posture information based on a visual configuration interface, and encodes the user-defined information into structured character feature parameters. The product images are subjected to noise reduction, enhancement, and contour extraction operations to extract the basic features of product color, texture, and shape, forming standardized product information data; Based on the user's defined target audience, marketing scenarios, and promotional information, the data is mapped to marketing configuration parameters that can be recognized by the algorithm, resulting in customized demand data with a unified format.
[0012] In one implementation of this application, after generating the preliminary synthesized video, the method further includes: In response to enterprise-level customization instructions, brand customization elements are modularly integrated into the initial synthesized video to generate the final marketing short video; the brand customization elements include brand logo images, brand-specific color cards, and brand scene images.
[0013] In one implementation of this application, in response to enterprise-level customization instructions, brand customization elements are modularly integrated into the edited video content to generate the final marketing short video, specifically including: The system receives brand-customized elements uploaded by users and performs real-time semantic analysis on each frame of the video content to identify core content areas and safe areas in the frame, generating regional semantic analysis data. The brand-customized elements include brand logo images, brand-specific color cards, and brand scene images. Based on the semantic analysis data of the region, the implantation position, scaling ratio and transparency parameters of the brand logo image within the safe area are dynamically calculated to generate brand logo implantation parameters; Based on the brand-specific color chart, the overall color tone of the video content is adjusted to generate color-adapted video frames. Based on the brand scene image, the brand scene is replaced or superimposed with the original background in the video to generate a video frame after scene fusion. Based on the brand logo embedding parameters, the video frames after color adaptation, and the video frames after scene fusion, the preliminary synthesized video is modified frame by frame to complete the integration of brand-customized elements.
[0014] On the other hand, embodiments of this application also provide an artificial intelligence-based product marketing short video generation device, the device comprising: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are then executed by the at least one processor to enable the at least one processor to perform the AI-based product marketing short video generation method described above.
[0015] On the other hand, this application embodiment also provides a non-volatile computer storage medium storing computer-executable instructions, which, when executed, implement the above-described method for generating short product marketing videos based on artificial intelligence.
[0016] This application provides a method and device for generating short product marketing videos based on artificial intelligence, which has at least the following beneficial effects: By establishing a three-dimensional collaborative mechanism between people, products, and marketing, and performing weighted fusion calculations based on product attribute tags, marketing scenario tags, and target audience profiles, precise matching of persona style and action sequence was achieved, enhancing the marketing adaptability and persuasiveness of video content. By constructing an intelligent recognition mechanism and scenario-based display scheme for the core selling points of products, the identified core selling point areas were integrated with a pre-built scenario material library to generate a display scheme that naturally combines the product with the usage scenario. Furthermore, the display order and duration were automatically planned according to marketing logic, resulting in a substantial improvement in the accuracy and attractiveness of product display. By establishing a collaborative matching mechanism for customized audio and video shot sequences, the problem of independent audio and video content and poor marketing synergy in existing solutions was solved. Finally, by introducing a modular integration mechanism of pre-set timing templates and enterprise-level customized elements, a complete marketing loop from product introduction to purchase guidance was achieved, meeting the enterprise-level deep customization needs for brand differentiation. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating the AI-based product marketing short video generation method provided in this application embodiment; Figure 2 A schematic diagram of the internal structure of an AI-based product marketing short video generation device provided in this application embodiment. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0020] Figure 1 This is a flowchart illustrating the AI-based product marketing short video generation method provided in this application embodiment.
[0021] The analysis method involved in the embodiments of this application can be implemented by a terminal device or a server, and this application does not impose any special limitations on it. For ease of understanding and description, the following embodiments are all described in detail using a server as an example.
[0022] It should be noted that the server can be a single device or a system composed of multiple devices, i.e., a distributed server. This application does not make any specific limitations on this.
[0023] like Figure 1 As shown in the embodiments of this application, the method for generating short product marketing videos based on artificial intelligence includes: Step 101: Receive and standardize the customized demand data input by the user, and based on the processed customized demand data, perform a three-dimensional weighted fusion calculation of the matching degree of audience preferences by matching the basic association weight of product attribute tags, matching the scene adaptation coefficient of marketing scene tags, and calculating the audience preference matching degree of the target audience profile, and generate an adapted person display scheme.
[0024] It should be noted that the customized requirement data in this application embodiment includes personal characteristic parameters, product information data, and marketing configuration parameters.
[0025] In this embodiment, the system provides a visual configuration interface, through which users can upload custom character image files, such as 3D models or 2D image sequences, or select from a variety of preset basic image libraries. Simultaneously, users can set the character's gender, age range, style type, and desired poses via drop-down menus or slider controls. The system uniformly encodes these scattered input information into structured character feature parameters and stores them in an internally common data format, such as using a dictionary structure to record fields like image identifiers, style codes, and action codes. It should be noted that for product information data, after receiving the product images uploaded by the user, the system first performs image preprocessing operations, including noise reduction filtering to remove acquisition noise, contrast enhancement to highlight product details, and contour extraction to obtain the main boundaries of the product.
[0026] The preprocessed output of basic product features such as color distribution, texture characteristics, and shape edges, together with the product attribute tags directly filled in by the user, such as category, key features, and selling point text, constitutes standardized product information data. It can be understood that the standardization of marketing configuration parameters involves mapping the user-defined target audience, such as age range, gender preference, and consumption preferences, marketing scenarios such as Douyin live streaming traffic generation and Taobao product promotion, and promotional information such as limited-time discounts and spending-based reduction activities, to field values in a pre-defined marketing parameter template, thereby forming a uniformly formatted marketing configuration parameter.
[0027] In this embodiment, the system first obtains the basic association weights for matching from a preset adaptation rule base based on the product attribute tags in the product information data. This adaptation rule base is trained using machine learning algorithms based on massive amounts of product marketing case data. It records the association weight matrix between different product types and character styles and gestures. For example, beauty products have a high basic association weight with professional review styles or lifestyle recommendation styles, while home appliance products have a high basic association weight with functional demonstration styles. It should be noted that the basic association weights reflect the general compatibility between the product type and a particular character style combination under no other constraints.
[0028] Next, the system retrieves the matching scenario adaptation coefficient from the adaptation rule library based on the marketing scenario tags in the marketing configuration parameters. For example, a live streaming traffic generation scenario on one platform requires fast-paced, visually impactful content, and the corresponding scenario adaptation coefficient will increase the weight of the person's style, which is lively and speaks quickly; while a product detail page on another platform focuses more on displaying product details, and the corresponding coefficient will increase the weight of the person's style, which is composed and features many close-up shots.
[0029] Simultaneously, the system calculates the audience preference matching degree from the adaptation rule base based on the target audience profile in the marketing configuration parameters. It's understandable that audience groups of different ages, genders, and consumption preferences have different tastes in character styles. For example, younger audiences tend to prefer trendy and fun character styles, while middle-aged and older audiences prefer stable and trustworthy styles. The system calculates the audience preference matching degree by comparing the target audience profile with the audience acceptance data for each style stored in the rule base. Subsequently, the system performs a weighted fusion calculation of the basic association weight, scenario adaptation coefficient, and audience preference matching degree to generate the final adaptation weight. This fusion calculation can use linear weighting or a non-linear combination method, and the weights of each coefficient can be dynamically adjusted according to the actual application scenario.
[0030] Finally, based on the final ranking results of the adaptation weights, the system selects the optimal character style type and action posture sequence from the adaptation rule base, and combines this information to generate an adapted character display scheme. This character display scheme not only includes the rendering parameters of the character image, but also the temporal description of the actions and postures, such as what actions to perform in the video introduction stage and what actions to perform in the selling point explanation stage, thus providing a basis for subsequent video arrangement. This achieves multi-dimensional and accurate matching between character characteristics and product attributes, marketing scenarios, and audience preferences, solving the technical problem of mismatch between character style and product selling point communication needs in existing solutions.
[0031] In one embodiment of this application, to address the shortcomings of traditional collaborative filtering algorithms, a three-dimensional weight factor is introduced to construct the adaptation weight calculation formula as follows: ,in, Indicates the basic association weight. Indicates the scene adaptation coefficient. This indicates the audience preference matching degree. Specifically, the adaptation rule base is trained using an embedding model based on a graph neural network. The training dataset contains manually labeled, compliant product marketing short video cases, covering 30 major product categories such as beauty, home appliances, food, and maternal and infant products. For each case data, its product attribute tags (such as category, price range, core functions), the type of persona used (such as professional review, lifestyle sharing, fun product recommendations), action posture sequences (such as handheld display, operation demonstration, comparative experiment), and the corresponding marketing scenario (such as live streaming traffic generation, product detail page display) and target audience profile (such as age range, gender distribution, and consumption level) are extracted. The graph neural network model uses product attribute tags and persona style types as nodes, and co-occurrence frequency as edge weights, learning the high-dimensional embedding representation of nodes through multi-layer graph convolution operations.
[0032] It should be noted that the basic association weight The calculation method is as follows: for the input product attribute tag, its embedding vector is queried in the graph neural network, the cosine similarity with the embedding vectors of various character styles is calculated, and the initial weight distribution is obtained after Softmax normalization. Scene adaptation coefficient The determination method involves pre-defining a scenario-style adaptation matrix. Matrix elements represent the gain coefficient of a certain persona style in a specific marketing scenario. For example, in a live-streaming traffic generation scenario, the coefficient for a lively, fast-paced style is increased to 1.3; in a product details page scenario, the coefficient for a composed, close-up style is increased to 1.2. This matrix is obtained through statistical analysis of historical case data and is dynamically updated according to the platform's algorithm recommendation rules. Audience preference matching degree. The calculation method involves encoding the target audience profile (age, gender, consumption preferences) into a multi-dimensional vector, and then performing KL divergence calculations with the audience acceptance distributions for each style stored in the matching rule base. A smaller divergence value indicates a higher degree of matching. After normalization, the result is... Value. Final adaptation weight. The product of the three factors is used. For example, the system selects the top three character style types and action posture sequences by weight as candidate solutions for user confirmation or automatic selection of the optimal solution.
[0033] Step 102: Based on product information data and marketing configuration parameters, identify and extract the core selling points visual areas of the product, and generate a product display scheme that integrates the product and usage scenarios based on the identification results and the pre-set scene material library.
[0034] In this embodiment, the system first preprocesses the product images in the product information data to extract multi-scale basic image features. For example, through different layers of a convolutional neural network, the system can acquire multi-level visual representations ranging from low-level features such as edges and corners to high-level features such as components and structures. Simultaneously, the system performs semantic encoding on the product attribute tags in the product information data, converting discrete tag text into semantic vectors in a continuous vector space. For example, for product attribute tags such as lactose-free, high-calcium, and fat-free, the semantic encoding model maps them into vector representations reflecting nutritional and health dimensions. It is understood that the semantic encoding model can be fine-tuned using a pre-trained language model to accurately capture semantic information in the product marketing domain.
[0035] For example, the product image is first preprocessed, using ResNet-50 as the backbone network to extract multi-scale image features, with an output feature map size of 14×14×2048. Simultaneously, the product attribute labels are semantically encoded, using a finely tuned BERT-base model to convert the label text into a 768-dimensional semantic vector, which is then mapped to 2048 dimensions through a fully connected layer to match the channel dimensions of the image features.
[0036] Next, the system calculates semantic attention weights for the basic image features based on the product attribute semantic vector. It's important to note that this semantic attention weight calculation process aims to focus the model on image regions strongly related to the product's selling points. Specifically, the system concatenates or multiplies the semantic vector with the image feature map, and a lightweight attention generation network outputs an attention weight map of the same size as the feature map. Regions with higher values in the attention weight map correspond to image locations semantically related to the product's selling points. For example, for a dog food product emphasizing high fresh meat content, the attention weight map will focus on the text area printed on the product packaging indicating the fresh meat content, rather than simply the overall outline of the dog food bag.
[0037] For example, the semantic vector and the image feature map are concatenated along the channel dimension and input into a lightweight attention generation network consisting of two convolutional layers and a sigmoid activation function. The output is a single-channel attention weight map with a size of 14×14, where each pixel value represents the association strength between the corresponding image region and the semantic meaning of the product's selling points. The convolutional kernel sizes are 3×3 and 1×1, with 512 and 1 channels, respectively. This attention weight map is then element-wise multiplied with the original image feature map to achieve feature recalibration, allowing the model to focus on regions related to the semantic meaning of the selling points.
[0038] Subsequently, the system drives the object detection model to identify product images based on the attention weight feature map. For example, the object detection model can employ an improved YOLO series or Faster R-CNN architecture. The core improvement lies in using the attention weight map as spatial guidance, focusing the detection model's region of interest on areas with high attention weights, rather than searching uniformly across the entire image. Finally, the detection model outputs core selling point visual region data, which includes the bounding box coordinates, category labels, and confidence scores for each selling point region. Category labels include, for example, nutrition facts, product logo, and usage effect images. This achieves visual focus from general object detection to marketing semantic guidance, improving the accuracy of selling point region identification.
[0039] For example, based on the recalibrated feature map, YOLOv8-nano is used as the target detection model to identify the core selling point region, and the bounding box coordinates, class label and confidence score are output. The confidence threshold is set to 0.75 and the non-maximum suppression IoU threshold is set to 0.5.
[0040] In this embodiment, the system retrieves corresponding background scene images from a pre-set scene material library based on the usage scenario tags in the product attribute tags. For example, outdoor equipment products correspond to natural landscape scenes, kitchen appliances to modern kitchen scenes, and pet food to cozy living room scenes. It is understood that the scene material library stores background images of various styles and lighting conditions, along with semantic description tags for the scenes, to facilitate accurate matching based on product attributes.
[0041] During the fusion process, the system employs an image fusion algorithm to synthesize the extracted core selling point areas with the background scene image. It's worth noting that this fusion algorithm goes beyond simple image overlay; it also considers perspective and the consistency of lighting effects, ensuring the product appears naturally placed within the scene. For example, when a dog food bag is placed on a living room coffee table, the algorithm generates appropriate shadows for the bag based on the direction of light in the background scene; when a milk carton is placed on a breakfast table, the algorithm adjusts the milk carton's color saturation to match the overall color tone of the scene.
[0042] After completing the scenario-based presentation, the system also needs to plan the pace of the product display based on the marketing logic in the marketing configuration parameters. For example, the system determines the order, duration, and camera movement of each core selling point area in the video. This might involve first showing the overall product appearance, then taking close-ups of each selling point area, and finally demonstrating the product's integration with the scenario. This solves the technical problems of vague product presentation and unclear selling points in existing solutions, making product presentations more precise and attractive.
[0043] Step 103: Based on the character display plan, product display plan, product information data, and marketing configuration parameters, generate customized audio that matches the rhythm of the video content and the marketing logic, and establish a collaborative matching relationship between the customized audio and the video shot sequence.
[0044] In this embodiment, the system first generates a segmented script based on the product selling point text in the product information data, the marketing scenario rules in the marketing configuration parameters, and the target audience preferences, using a pre-trained language model. It should be noted that this language model is fine-tuned to adapt to the language style of the marketing scenario, enabling it to divide the entire script into several segments according to marketing logic, with each segment corresponding to a marketing node. For example, the generated script segments include a product introduction segment, a selling point breakdown segment, a trust-building segment, and a purchase guidance segment. Each segment contains specific explanatory phrases, and the model generates a timestamp for each script segment to indicate its approximate time position within the overall audio. It is understood that segmented script generation is for subsequent audio and video alignment.
[0045] Next, based on the segmented speech scripts and the character style types in the character display scheme, the system calls the corresponding emotional speech synthesis model. Specifically, the emotional speech synthesis model can adjust the timbre, speech rate, pitch, and emotional color of the synthesized speech according to the style tags. For example, a gentle and lifelike style corresponds to a moderate speech rate and a soft timbre, an energetic and trendy style corresponds to a faster speech rate and a bright timbre, and a professional evaluation style corresponds to a stable speech rate and a neutral timbre. The speech synthesis model takes the segmented speech scripts as input and outputs segmented speech audio files with precise timestamps.
[0046] Simultaneously, the system also needs to match background music. For example, based on the marketing scenario rules in the marketing configuration parameters and the overall rhythm of the segmented voice script, the system retrieves suitable music clips from the background music library. For instance, promotional scenarios require fast-paced, energetic music, while brand promotion scenarios require soothing, grand music. After matching, the system performs volume balancing and mixing processing on the segmented voice audio and background music clips to ensure clear and indistinguishable speech without the background music overpowering the main audio, ultimately generating complete customized audio data.
[0047] In this embodiment, the system first parses the timestamps in the customized audio to obtain audio content segment nodes, i.e., the start and end times of each script segment. Simultaneously, the system plans the video shot sequence according to a timing arrangement template to obtain video shot segment nodes. It should be noted that the timing arrangement template generates a synchronized audio-visual data stream with fixed time window constraints by aligning the audio content segment nodes with the video shot segment nodes using timestamps. The video shot sequence originates from character display schemes and product display schemes, which include preset durations and content types for each shot.
[0048] The system aligns audio content segments with video shot segments based on preset mapping rules. For example, it aligns product introduction segments with panoramic shots of people and products, and selling point explanation segments with close-up shots of selling points, thus generating an initial audio-visual synchronization alignment scheme. However, since there may be slight differences between the duration of the synthesized speech and the preset shot duration, the system also needs to perform real-time comparison and dynamic adjustment. Specifically, the system compares the duration of each segment of the audio content segments with the corresponding duration of the video shot segments in the initial alignment scheme, calculating the duration deviation between the two. When the duration deviation exceeds a preset threshold, the system triggers a dynamic adjustment mechanism: if the audio duration is longer than the shot duration, the system can slightly increase the shot playback rate or slightly extend the end of the shot; if the audio duration is shorter than the shot duration, the system can slightly decrease the shot playback rate or make minor adjustments to the speech rate. Through this closed-loop adjustment, the system generates an adjusted audio-visual synchronization alignment scheme, achieving high-precision audio-visual synchronization. Understandably, this collaborative matching mechanism ensures that the product selling points in the voice explanation and the character's operation demonstration in the video correspond precisely in time, thereby improving the coherence of the video content and the persuasiveness of the marketing.
[0049] For example, when the absolute value of the deviation exceeds a threshold of 0.3 seconds, an adjustment strategy is triggered. If the audio duration is longer than the shot duration, a linear interpolation algorithm is used to increase the shot playback rate by 5% to 10%, or a frame repetition algorithm is used to add transition frames at the end of the shot. If the audio duration is shorter than the shot duration, a Waveform Similarity Overlap-Add (WSOLA) algorithm is used to extend the duration of the speech, with the extension ratio controlled within 8% to maintain the naturalness of the speech, or a shot rate reduction strategy is used, so that the adjusted audio-visual synchronization deviation can be stably controlled within 0.2 seconds according to actual tests.
[0050] Step 104: Based on the preset timing arrangement template, automatically arrange and integrate the character display scheme, product display scheme and customized audio to generate a preliminary synthesized video.
[0051] In this embodiment, the system first acquires character motion rendering data from the character display scheme, product and scene image sequences from the product display scheme, and customized audio data. Each of these data carries both time and content attributes. The system internally uses a pre-defined time-series orchestration template, which addresses the synchronization and integration of multimodal data streams (audio streams, video streams, and image sequences) over time. It should be noted that the time-series orchestration template is a structured configuration file defining the stage order, duration range, and content type constraints. This template includes four segments: a first time-series segment, a second time-series segment, a third time-series segment, and a fourth time-series segment. Each segment corresponds to a specific combination of content types and duration allocation suggestions. The time-series orchestration template is stored in the system configuration library in JSON format. The duration allocation and content mapping rules for each stage can be quickly adjusted by modifying the configuration file without modifying the core algorithm code. The system maps and associates the character display scheme, product display scheme, and customized audio with the various stages in the pre-defined time-series orchestration template based on their content attributes, generating the original content mapping relationship.
[0052] For example, the system calls the character rendering engine to generate a sequence of character action frames, while overlaying the overall product image. An audio stream plays a character introduction voice, mapping the panoramic shot of the character and the product introduction dialogue to the first time segment, with a preset duration of 3 seconds. The system generates a sequence of close-up shots based on the visual area data of the core selling points, achieving natural transitions between frames through a smooth transition algorithm (such as optical flow). An audio stream plays a voice explaining the selling points, mapping the close-up shots and explanations to the second time segment, with the duration evenly distributed according to the number of selling points. The system calls preset usage scenario video clips or generates animation frames simulating usage effects. An audio stream plays a trust endorsement voice, mapping the usage effect demonstration and user reviews to the third time segment, with a preset duration of 5 seconds. The system renders a graphic layer containing price tags and countdown components and overlays it onto the video frames. An audio stream plays guiding voice, mapping the display of promotional information and click indicators to the fourth time segment, with a preset duration of 10 seconds.
[0053] Next, based on the target audience and promotional information in the marketing configuration parameters, the system personalizes the content mapping relationship by filling in content and allocating duration accordingly. It's important to note that the emphasis of the marketing logic differs for different target audiences. For example, for younger audiences, the system might add fast-paced transitions and fashionable elements, emphasizing the urgency of limited-time offers during the purchase guidance phase; for middle-aged and older audiences, the system might slow down the pace, increase the duration of trust-endorsing content, and emphasize safety and reliability during the purchase guidance phase. After completing the arrangement, the system segments and aligns the customized audio according to the video content structure, ensuring accurate alignment of audio segments and corresponding video segments on the timeline. Simultaneously, the system renders the character presentation and product presentation schemes as video streams, integrates them according to the video content structure, and generates a preliminary composite video.
[0054] In this embodiment, the system receives brand-customized elements uploaded by users. These elements reflect the company's brand image and differentiated marketing needs. It should be noted that brand-customized elements include brand logo images, brand-specific color swatches, and brand scene images. Brand logo images include logo files, brand-specific color swatches include the RGB values of the primary and secondary colors, and brand scene images include photos of brand stores or virtual backgrounds. The system performs real-time semantic analysis on each frame of the initially synthesized video. Using a lightweight convolutional neural network, it identifies core content areas such as faces, product subjects, key text areas, and safe areas with minimal visual interference, such as solid-color background areas in the corners of the frame, thereby generating regional semantic analysis data.
[0055] Based on this semantic analysis data, the system dynamically calculates embedding parameters for the brand logo image. Specifically, it determines the logo's scaling ratio based on the size of the safe area, ensuring it doesn't obscure the core content while maintaining recognizability; it determines the transparency based on the brightness and texture complexity of the safe area, allowing the logo to blend naturally with the background rather than appearing abruptly pasted; and it determines the embedding coordinates based on the location of the safe area. For example, the system uses Poisson blending technology to embed the logo into a specified position in the video frame, achieving a natural edge transition. Simultaneously, the system performs color matching adjustments to the overall tone of the video content based on a brand-specific color swatch, such as mapping the video's global tone to near the brand's primary color, or overlaying the brand's color swatch in transition shots as a transition effect.
[0056] Furthermore, for brand scene images, the system replaces or overlays them with the original background in the video. For example, it uses photos of brand stores as the background for interviews or virtual brand avatar scenes as the background for product displays, ensuring realism through perspective correction and lighting adaptation. After frame-by-frame modification of all customized elements, the system generates the final marketing short video. This short video can be directly used for live streaming on e-commerce platforms or for product promotion activities, possessing a complete marketing loop and a distinct brand personality.
[0057] Compared to existing technologies, this application achieves precise matching between personality style and product characteristics by constructing a three-dimensional weighted fusion calculation mechanism of personality feature parameters, product attribute tags and marketing scenario tags. By introducing attention weight calculation driven by product attribute semantic vectors, it achieves intelligent recognition and focused display of the visual area of the product's core selling points. By establishing an alignment and dynamic adjustment mechanism between audio content segment nodes and video shot segment nodes, it achieves high-precision synchronization of audio and video content at marketing logic nodes.
[0058] Specifically, under the same testing conditions, using images of pet food products as input, existing technologies can only identify the overall outline of the product and cannot locate the text area of the core selling point "70% fresh meat content." However, the semantic attention guidance mechanism of this application can accurately identify this area and assign it a high-weight attention value, increasing the recognition accuracy from approximately 65% in existing technologies to 92.3%. Regarding audio-visual synchronization, existing technologies lack a timing alignment mechanism, resulting in a timing deviation of more than 0.8 seconds between the voice explanation and the product demonstration. This application, through timestamp marking and dynamic rate adjustment, controls the deviation to within 0.2 seconds. In terms of the compatibility between people and products, existing technologies use fixed template matching with an accuracy rate of approximately 58%, while the three-dimensional weighted fusion computing mechanism of this application increases the accuracy rate to 92.5%.
[0059] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide an artificial intelligence-based product marketing short video generation device, the structure of which is as follows: Figure 2 As shown.
[0060] Figure 2 This is a schematic diagram of the internal structure of an AI-based product marketing short video generation device provided in an embodiment of this application. Figure 2 As shown, the device includes: At least one processor; And, a memory that is communicatively connected to at least one processor; The memory stores instructions that can be executed by at least one processor, and the instructions, when executed by at least one processor, enable at least one processor to: The system receives and standardizes customized user input data, and based on the processed customized data, performs a three-dimensional weighted fusion calculation of product attribute tag matching basic association weight, marketing scenario tag matching scenario adaptation coefficient, and target audience profile to calculate audience preference matching degree. It then performs marketing adaptation processing between people and products to generate an adapted person display scheme. The customized data includes person feature parameters, product information data, and marketing configuration parameters. Based on product information data and marketing configuration parameters, identify and extract the core selling points visual areas of the product, and generate a product display solution that integrates the product with the usage scenario based on the identification results and a pre-set scene material library. Based on the character display plan, product display plan, product information data, and marketing configuration parameters, a customized audio is generated that matches the rhythm of the video content and the marketing logic, and a collaborative matching relationship is established between the customized audio and the video shot sequence. Based on preset timing templates, the system automatically arranges and integrates character presentation schemes, product presentation schemes, and customized audio to generate a preliminary composite video.
[0061] This application also provides a non-volatile computer storage medium storing computer-executable instructions, which, when executed, can: The system receives and standardizes customized user input data, and based on the processed customized data, performs a three-dimensional weighted fusion calculation of product attribute tag matching basic association weight, marketing scenario tag matching scenario adaptation coefficient, and target audience profile to calculate audience preference matching degree. It then performs marketing adaptation processing between people and products to generate an adapted person display scheme. The customized data includes person feature parameters, product information data, and marketing configuration parameters. Based on product information data and marketing configuration parameters, identify and extract the core selling points visual areas of the product, and generate a product display solution that integrates the product with the usage scenario based on the identification results and a pre-set scene material library. Based on the character display plan, product display plan, product information data, and marketing configuration parameters, a customized audio is generated that matches the rhythm of the video content and the marketing logic, and a collaborative matching relationship is established between the customized audio and the video shot sequence. Based on preset timing templates, the system automatically arranges and integrates character presentation schemes, product presentation schemes, and customized audio to generate a preliminary composite video.
[0062] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0063] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0069] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0070] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0071] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes said element.
[0072] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating a short video for product marketing based on artificial intelligence, characterized in that, The method includes: The system receives and standardizes user-inputted customized demand data. Based on the processed customized demand data, it performs a three-dimensional weighted fusion calculation of product attribute tag matching basic association weight, marketing scenario tag matching scenario adaptation coefficient, and target audience profile to calculate audience preference matching degree. It then performs person-product marketing adaptation processing to generate an adapted person display scheme. The customized demand data includes person feature parameters, product information data, and marketing configuration parameters. Based on the product information data and the marketing configuration parameters, the core selling points visual areas of the product are identified and extracted. Based on the identification results and the pre-set scene material library, a product display scheme that integrates the product and the usage scenario is generated. Based on the character display scheme, the product display scheme, the product information data, and the marketing configuration parameters, a customized audio is generated that matches the rhythm of the video content and the marketing logic, and a collaborative matching relationship is established between the customized audio and the video shot sequence. Based on a preset time-series arrangement template, the character display scheme, the product display scheme, and the customized audio are automatically arranged and integrated to generate a preliminary synthesized video. 2.The AI-based product marketing short video generation method of claim 1, wherein, Based on the processed customized demand data, a three-dimensional weighted fusion calculation is performed using product attribute tag matching basic association weights, marketing scenario tag matching scenario adaptation coefficients, and target audience profile to calculate audience preference matching degree. This process then executes persona-product marketing adaptation processing to generate an adapted persona display scheme, specifically including: Based on product attribute tags from product information data, the basic association weights for matching are obtained from a preset adaptation rule base. Based on the marketing scenario tags in the marketing configuration parameters, the matching scenario adaptation coefficient is obtained from the adaptation rule base, and based on the target audience profile in the marketing configuration parameters, the audience preference matching degree is calculated from the adaptation rule base. The basic association weight, the scene adaptation coefficient and the audience preference matching degree are weighted and fused to generate the final adaptation weight; Based on the ranking results of the final adaptation weights, the optimal character style type and action posture sequence are selected from the adaptation rule base and combined to generate an adapted character display scheme. 3.The AI-based product marketing short video generation method of claim 1, wherein, Based on the product information data and the marketing configuration parameters, the core selling points visual areas of the product are identified and extracted, specifically including: The product images in the product information data are preprocessed to extract multi-scale basic image features, and the product attribute tags in the product information data are semantically encoded to generate product attribute semantic vectors. Based on the product attribute semantic vector, semantic attention weights are calculated on the basic features of the image to generate an attention weight feature map that focuses on the semantically related areas of the product's selling points. Based on the attention weight feature map, the target detection model is driven to identify product images and output core selling point visual region data; the core selling point visual region data includes the coordinates, category and confidence level of the core selling point region. 4.The AI-based product marketing short video generation method of claim 1, wherein, Based on the aforementioned character display scheme, product display scheme, product information data, and marketing configuration parameters, customized audio is generated to match the rhythm of the video content and the marketing logic, specifically including: Based on the product selling point text in the product information data, the marketing scenario rules in the marketing configuration parameters, and the target audience preferences, a segmented script is generated using a pre-trained language model, and a segmented speech script with a marketing logic structure is output. Based on the segmented speech script and the character style type in the character display scheme, the corresponding emotional speech synthesis model is called to generate segmented speech audio with timestamps. Based on the marketing scenario rules in the marketing configuration parameters and the overall rhythm of the segmented voice script, the corresponding background music segments are matched and extracted from the background music library; The segmented audio and background music clips are then mixed and balanced in volume to generate the final customized audio. 5.The AI-based product marketing short video generation method of claim 4, wherein, Establishing the collaborative matching relationship between the customized audio and video shot sequences specifically includes: The timestamp markers in the customized audio are parsed to obtain audio content segmentation nodes, and the video shot sequence is planned according to the time sequence arrangement template to obtain video shot segmentation nodes; Based on the preset mapping rules, the audio content segment nodes are aligned and matched with the video shot segment nodes to generate an initial audio-visual synchronization alignment scheme. The duration of each segment of the audio content segment node in the initial audio-visual synchronization alignment scheme is compared with the duration of the corresponding segment of the video shot segment node in real time, and the duration deviation is calculated. When the duration deviation exceeds a preset threshold, the playback rate of the video shot sequence or the speech rate of the customized audio is dynamically adjusted to generate an adjusted audio-visual synchronization alignment scheme to achieve audio-visual synchronization. 6.The AI-based product marketing short video generation method of claim 1, wherein, Based on a preset time-series arrangement template, the character display scheme, the product display scheme, and the customized audio are automatically arranged and integrated to generate a preliminary synthesized video, specifically including: The character display scheme, the product display scheme, and the customized audio are mapped and associated with each stage in the preset time sequence arrangement template according to the content attributes to generate the original content mapping relationship. The time sequence arrangement template is a structured configuration file that defines the stage order, duration range, and content type constraints. It includes a first time sequence segment, a second time sequence segment, a third time sequence segment, and a fourth time sequence segment. Each segment corresponds to different content type combinations and different time windows with durations greater than a threshold. Based on the target audience and promotional information in the marketing configuration parameters, the original content mapping relationship data is personalized with content filling and duration allocation to generate the arranged video content structure; The customized audio is segmented and aligned according to the video content structure, and the character display scheme and the product display scheme are rendered into video streams. They are then integrated according to the video content structure to generate a preliminary composite video. 7.The AI-based product marketing short video generation method of claim 1, wherein, Receive and standardize user-input customized request data, specifically including: The system collects user-defined character image, gender, age, style, and action posture information based on a visual configuration interface, and encodes the user-defined information into structured character feature parameters. The product images are subjected to noise reduction, enhancement, and contour extraction operations to extract the basic features of product color, texture, and shape, forming standardized product information data; Based on the user's defined target audience, marketing scenarios, and promotional information, the data is mapped to marketing configuration parameters that can be recognized by the algorithm, resulting in customized demand data with a unified format. 8.The AI-based product marketing short video generation method of claim 1, wherein, After generating the initial synthesized video, the method further includes: In response to enterprise-level customization instructions, brand customization elements are modularly integrated into the initial synthesized video to generate the final marketing short video; the brand customization elements include brand logo images, brand-specific color cards, and brand scene images. 9.The AI-based product marketing short video generation method of claim 8, wherein, Responding to enterprise-level customization instructions, brand customization elements are modularly integrated into the initial composite video to generate the final marketing short video, specifically including: The system receives brand-customized elements uploaded by users and performs real-time semantic analysis on each frame of the video content to identify core content areas and safe areas in the frame, generating regional semantic analysis data. The brand-customized elements include brand logo images, brand-specific color cards, and brand scene images. Based on the semantic analysis data of the region, the implantation position, scaling ratio and transparency parameters of the brand logo image within the safe area are dynamically calculated to generate brand logo implantation parameters; Based on the brand-specific color chart, the overall color tone of the video content is adjusted to generate color-adapted video frames. Based on the brand scene image, the brand scene is replaced or superimposed with the original background in the video to generate a video frame after scene fusion. Based on the brand logo embedding parameters, the video frames after color adaptation, and the video frames after scene fusion, the preliminary synthesized video is modified frame by frame to complete the integration of brand-customized elements.
10. An apparatus for generating a short video for product marketing based on artificial intelligence, characterized by, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the AI-based product marketing short video generation method as described in any one of claims 1-9.
Citation Information
Patent Citations
E-commerce advertisement video script dynamic generation method and system based on data driving
CN121691849A