Marketing video generation method and system, storage medium and electronic equipment

By obtaining production instructions for marketing products and using a dual-channel parallel video generation system to generate product close-ups and scene mood sequences, the problem of monotonous marketing video display effects is solved, personalized marketing video generation is realized, and the display effect is improved.

CN121078291APending Publication Date: 2025-12-05SUZHOU YIMAI DONGXI INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511320018.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing marketing video production techniques mainly use fixed templates, resulting in monotonous marketing video presentations and poor product marketing effectiveness.

Method used

By obtaining the production instructions for the target marketing product, including product materials, usage scenarios, target audience and additional needs, the key features of the product are determined. A dual-path parallel video generation system is used to generate product close-up sequences and scene mood sequences respectively. The video is then segmented to generate a visual element library. Narrative scripts and audio scripts are generated based on the group characteristics of the target audience. Finally, the marketing video is generated by combining the sequences in time.

Benefits of technology

It enables personalized arrangement of marketing videos, highlights product features and creates suitable scene atmosphere, thereby improving the display effect of marketing videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078291A_ABST
    Figure CN121078291A_ABST
Patent Text Reader

Abstract

The invention discloses a marketing video generation method and system, a storage medium and electronic equipment, and relates to the technical field of video processing. The method comprises the steps of obtaining a manufacturing instruction of a target marketing product, and determining key features of the product according to additional requirements and product materials; generating a product close-up sequence based on the product materials and the product key features, and generating a scene artistic conception sequence based on the use scene and the product key features in the second path; performing picture segmentation on the product close-up sequence and the scene artistic conception sequence to generate a visual element library; generating a narrative script and an audio script according to the group characteristics of the target audience group and the product key characteristics, arranging the visual element library according to the narrative script, and generating a picture display sequence; generating an audio sequence corresponding to the picture display sequence according to the audio script; and performing time sequence combination on the picture display sequence and the audio sequence to generate a marketing video of the target marketing product. By implementing the technical scheme provided by the invention, the display effect of the marketing video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a marketing video generation method and system, a storage medium and an electronic device. BACKGROUND

[0002] With the development of digital marketing, video has become an important means of product display and brand promotion. Compared with traditional graphic marketing methods, video marketing can fully display product features through the combination of dynamic pictures and audio, effectively improving user understanding and cognition of products.

[0003] The existing marketing video production technology mainly adopts a material splicing method to generate a video. This technology generates a marketing video by combining product materials according to a preset template. In the specific implementation process, the system extracts product pictures, text descriptions and other materials according to fixed template rules, and arranges and combines these materials according to a preset layout. However, this template-based production method is too fixed, resulting in a single display effect of the generated marketing video and poor marketing effect of the product. SUMMARY

[0004] The present application provides a marketing video generation method, system, storage medium and electronic device, which can improve the display effect of the marketing video.

[0005] In a first aspect, the present application provides a marketing video generation method, which comprises: obtaining production instructions of a target marketing product, the production instructions at least including product materials, use scenarios, target audience groups and additional requirements; determining product key features according to the additional requirements and the product materials; loading a double-path parallel picture generation system including a first path and a second path, wherein the first path generates a product close-up sequence based on the product materials and the product key features, and the second path generates a scene mood sequence based on the use scenarios and the product key features; performing picture segmentation on the product close-up sequence and the scene mood sequence to generate a visual element library; generating a narrative script and an audio script according to the group characteristics of the target audience groups and the product key features, and arranging the visual element library according to the narrative script to generate a picture display sequence; generating an audio sequence corresponding to the picture display sequence according to the audio script; performing time sequence combination on the picture display sequence and the audio sequence to generate a marketing video of the target marketing product, and displaying the marketing video.

[0006] By adopting the technical scheme, the production instruction containing product material, use scene, target audience group and additional demand is acquired, the product key feature is determined based on the additional demand and the product material, then the product close-up sequence and the scene mood sequence are respectively generated by adopting the double-path parallel picture generation system, so that the marketing video can highlight the product features and create a suitable scene atmosphere; meanwhile, the visual element library is generated by picture segmentation on the product close-up sequence and the scene mood sequence, and the narration script and the audio script are generated according to the group features of the target audience group and the product key feature, so that the personalized arrangement of audio-visual content can be realized; finally, the picture display sequence arranged according to the narration script and the corresponding audio sequence are time-series combined, so that the marketing video highlighting the product features and having good audio-visual experience is generated, the technical problem that the display effect is single in the prior art due to the generation of the marketing video based on the fixed template is overcome, and the display effect of the marketing video is improved.

[0007] In a second aspect of the present application, a marketing video generation system is provided, the system comprising: a production instruction acquisition module configured to acquire production instructions of a target marketing product, the production instructions comprising at least product material, use scene, target audience group and additional demand; a key feature determination module configured to determine product key features according to the additional demand and the product material; a parallel sequence generation module configured to load a double-path parallel picture generation system comprising a first path and a second path, wherein the first path generates a product close-up sequence based on the product material and the product key features, and the second path generates a scene mood sequence based on the use scene and the product key features; an element library generation module configured to generate a visual element library by picture segmentation on the product close-up sequence and the scene mood sequence; a script generation module configured to generate a narration script and an audio script according to the group features of the target audience group and the product key features, and arrange the visual element library according to the narration script to generate a picture display sequence; an audio sequence generation module configured to generate an audio sequence corresponding to the picture display sequence according to the audio script; a marketing video generation module configured to time-series combine the picture display sequence and the audio sequence to generate a marketing video of the target marketing product, and display the marketing video.

[0008] In a third aspect of the present application, a computer storage medium is provided, the computer storage medium storing a plurality of instructions, the instructions being adapted to be loaded by a processor and execute the method steps described above.

[0009] In a fourth aspect of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded by the processor and execute the method steps described above.

[0010] In summary, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: The present application obtains production instructions containing product materials, use scenarios, target audience groups and additional requirements, determines product key features based on the additional requirements and product materials, and then generates product close-up sequences and scene mood sequences respectively using a double-path parallel picture generation system, so that the marketing video can highlight product features and create a suitable scene atmosphere. At the same time, by segmenting the product close-up sequences and scene mood sequences to generate a visual element library, and generating a narrative script and an audio script according to the group characteristics of the target audience groups and the product key features, personalized arrangement of audio-visual content can be achieved. Finally, the picture display sequence arranged according to the narrative script is combined with the corresponding audio sequence in time sequence, thereby generating a marketing video that can highlight product features and has good audio-visual experience, overcoming the technical problem of single display effect caused by generating marketing videos based on fixed templates in the prior art, and improving the display effect of marketing videos. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a flowchart of a marketing video generation method provided by an embodiment of the present application; Figure 2 is an exemplary user interaction interface diagram provided by an embodiment of the present application; Figure 3 is an exemplary marketing video display interface diagram provided by an embodiment of the present application; Figure 4 is another exemplary user interaction interface diagram provided by an embodiment of the present application; Figure 5 is another exemplary marketing video display interface diagram provided by an embodiment of the present application; Figure 6 is a module schematic diagram of a marketing video generation system provided by an embodiment of the present application; Figure 7 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0012] Legend of reference signs: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION EMBODIMENT

[0013] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in conjunction with the drawings in the specification embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, not all.

[0014] In the description of the embodiments of the present application, the words such as "for example" or "for instance" are used to represent an example, illustration or description. Any embodiment or design scheme described as "for example" or "for instance" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "for example" or "for instance" are intended to present the relevant concept in a specific manner.

[0015] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are used only for the purpose of description, and should not be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.

[0016] Please refer to Figure 1 , a flowchart of a marketing video generation method is specifically proposed, the method can be realized by relying on a computer program, can be realized by relying on a single-chip microcomputer, and can run on a marketing video generation system. The computer program can be integrated in a computer device, or can run as an independent tool application. Specifically, the method includes steps 10 to 70, and the above steps are as follows: Step 10: Obtain the production instruction of the target marketing product, and the production instruction at least includes product materials, use scenarios, target audience groups and additional requirements.

[0017] In the embodiments of the present application, the target marketing product refers to a specific product that needs to be promoted by video marketing. The product can be a physical commodity, a software application, a service project or other types of objects that can be commercially promoted.

[0018] The production instruction refers to a set of input information used to guide the generation of the marketing video, and specifically includes product materials (such as product pictures, videos, and text introductions, etc. basic data), usage scenarios (such as product application environment and usage situation, etc. scene information), target audience groups (such as target user age, occupation, and interest, etc. group attributes), and additional requirements (such as product features to be highlighted, marketing focus, brand tone, and other special requirements), which together constitute the basic input parameters of the video generation process.

[0019] Specifically, the system first needs to obtain the production instruction of the target marketing product, which serves as the basic input information for the entire video generation process. In specific implementation, the system receives user input information including but not limited to product materials, usage scenarios, target audience groups, and additional requirements, etc. through a user interface. Among them, the acquisition method of product materials includes: the user directly uploads product pictures, product video clips, product description documents, etc.; or the system automatically acquires the contents in the official material library of the product by calling the product database API (Application Programming Interface); or collects product-related materials from product websites, e-commerce platforms, etc. through crawler technology. For the acquisition of usage scenarios, the system provides a scene selection interface, including but not limited to office scenarios, home scenarios, outdoor scenarios, etc. preset options for users to choose from, while allowing users to input specific scene description texts. When obtaining target audience group information, the system collects attribute information such as target audience age, occupation type, consumption ability, and interest preference through a form, and can optionally associate user portrait information in historical marketing data. For the acquisition of additional requirements, the system provides a structured requirement collection form, including product selling points, marketing focus, brand tone, display style, and other dimensions of options, and users can specify the specific requirements through selection or text input. It should be noted that all the data obtained above are authorized.

[0020] After obtaining these production instructions, the system will preprocess and format the input information to ensure that the subsequent processing steps can correctly recognize and use these information. For example, file format verification and conversion of product materials, semantic standardization processing of scene descriptions, and data structuring of target audience information. This preprocessing mechanism can improve the stability and efficiency of the subsequent video generation process. At the same time, the system will also establish an association index of the input information to facilitate quick positioning and calling of related information in the subsequent feature extraction and content generation process.

[0021] Please refer to Figure 2 An exemplary user interaction interface diagram is provided for the embodiments of the present application, taking a smart earphone as an example, combined with Figure 2, the user gives the production instruction in the input mode, first the system gives the guide prompt word: "please enter the video requirements you want to make", then the user gives the production request based on the guide prompt word: "please generate a marketing video of a smart earphone, the product materials of the earphone are in the attachment, the theme should be simple and highlight the earphone appearance, and be used indoors", finally the system converts the production request into production instructions, that is, the user's production request is parsed into multiple instructions, such as the target product name, product related materials in the attachment, and additional requirements: the theme should be simple and highlight the earphone appearance, and other contents need to be parsed out.

[0022] Step 20: Determine the product key features according to the additional requirements and product materials.

[0023] In the embodiments of the present application, the product key features refer to the significant features that can highlight the core value of the product and distinguish it from other similar products, or the key features that the user wants to show.

[0024] In a feasible embodiment, a preset product feature category library can be established, including functional features, appearance features, and use experience features. After obtaining the additional requirements and product materials, the system extracts the feature information corresponding to the preset categories from the product materials through keyword matching. At the same time, the system retrieves the key display content specified in the additional requirements, and determines the features that appear in both the product materials and the additional requirements as key features. For example, when processing a certain smart watch, the system extracts the features "50 kinds of motion modes", "heart rate monitoring", and "15-day battery life" from the product materials, and finds that the "motion mode diversity" is also emphasized in the additional requirements, so "50 kinds of motion modes" is determined as a key feature.

[0025] In another feasible embodiment, the key features can also be determined based on the frequency statistics. The system performs word segmentation processing on the text content in the product materials, and counts the frequency of each word in the product manual, product description, and other materials. At the same time, the system also counts the number of times these words appear in the additional requirements. By setting a frequency threshold, the features whose total frequency in the product materials and the additional requirements exceeds the threshold are determined as key features. For example, when a certain feature word appears 5 times in the product materials and 3 times in the additional requirements, the total number of times exceeds the preset threshold of 7, then the feature is determined as a product key feature.

[0026] On the basis of the above embodiments, as another optional embodiment, the step of determining the product key features according to the additional requirements and product materials further includes the following steps: Step 101: Extract product basic information from product materials, including functional description information in product introduction text, appearance feature information in product images, and use effect information in product videos.

[0027] Specifically, product basic information is extracted from product materials through three dimensions. For function description information extraction: the system reads product manuals, product detail pages, and other text content, uses natural language processing technology for word segmentation and semantic extraction, and identifies all descriptions related to product functions, technical parameters, and performance indicators. For appearance feature extraction: the system analyzes product picture materials to extract product size data, material information, shape design, color scheme, and other static visual features. For use effect extraction: the system analyzes product demonstration videos to identify product use status, function display, and effect performance in different scenarios. For example, when processing a certain smart earphone, the system extracts "active noise reduction depth" and "battery life" from product documents, extracts "in-ear design" and "material technology" from product pictures, and extracts "noise reduction effect" and "anti-falling performance" from demonstration videos.

[0028] Step 102: Classify product basic information according to function dimension, appearance dimension, and use effect dimension to generate a feature candidate set.

[0029] Specifically, the system establishes a three-dimensional feature classification framework to structure and organize the extracted product basic information. In the function dimension, it is classified according to the technical field and application purpose of the function; in the appearance dimension, it is classified according to the design elements and visual features; and in the use effect dimension, it is classified according to the use scenario and experience effect. Through this classification method, all extracted feature information is organized into an ordered feature candidate set. For example, for a smart earphone product, the function features can be divided into core function class, audio performance class, and basic function class, the appearance features can be divided into structure design class and material technology class, and the use effect can be divided into different application scenario classes.

[0030] Step 103: Perform semantic analysis on additional requirements to identify user key display intentions.

[0031] Specifically, the additional demand text is structured for semantic analysis processing. First, keyword extraction is performed to identify key words and phrases in the text; then the occurrence frequency, position distribution and context relationship of these keywords are analyzed; finally, these information is integrated into specific display intent based on semantic understanding. This analysis method can accurately grasp the user's emphasis in the additional demand and the core direction of the expected display. For example, the production instruction is: "Please generate a marketing video for a smart earphone, the product materials of the earphone are in the attachment, the theme should be simple, and the focus should be on the earphone appearance, and it should be used indoors". The system determines the additional demand as: "the theme should be simple, and the focus should be on the earphone appearance, and it should be used indoors", and extracts the keywords "simple theme", "appearance", and "indoor" from the additional demand. Through semantic analysis, the system identifies the following display intent: visual style intent: simple style display; key content intent: highlight product appearance design; Scenario setting intent: indoor application scenario.

[0032] Step 104: Based on the user's key display intent, the importance of the features in the feature candidate set is scored, including display priority score, feature correlation score and expressiveness score.

[0033] Specifically, the system scores the importance of the features in the feature candidate set based on the identified user display intent in three dimensions: first, the display priority score is based on the matching degree between the feature and the user display intent. When the user display intent is "highlighting product appearance design", the features in the appearance dimension get higher priority scores. For example, "streamlined shell" gets 3 points, "matte texture" gets 3 points, and "noise reduction performance" gets only 1 point.

[0034] Second, the feature correlation score is based on the logical correlation strength between features. Since the user requires "simple style display", the system will analyze the correlation between features, and features that support each other get higher scores. For example, "streamlined shell" and "matte texture" form a simple design style, with high correlation between each other, each getting 2 points; while "multi-color color matching" does not match the simple style, with a low correlation score of only 0.5 points.

[0035] Finally, the expressiveness score is based on the showability of the feature in the video. Since the user specifies "indoor application scenario", the system needs to analyze the showability of the feature in the specified scenario. The system first performs brightness analysis on the feature region through image processing, calculates the average brightness value, the maximum and minimum brightness difference value, and the uniformity of the brightness distribution. Then, the feature clarity analysis is performed, the Sobel operator is used for edge detection to calculate the edge saliency score, and the texture complexity feature value is extracted through the gray level co-occurrence matrix. Finally, the spatial effect evaluation is performed to analyze the visible integrity, contour continuity and visual prominence of the feature under different viewing angles. Specifically, the system calculates the expressiveness score according to the following three dimensions: Brightness analysis score: 1 point for average brightness value in the range of 80-180, 1 point for maximum and minimum brightness difference value less than 100, 1 point for brightness distribution standard deviation less than 30, and the sum of the three scores constitutes the brightness analysis score; Clarity analysis score: 1 point for edge saliency greater than 0.6, 1 point for texture complexity in the range of 0.4-0.8, and the sum of the two scores constitutes the clarity analysis score; Spatial effect score: 1 point for visible integrity greater than 90%, 1 point for contour continuity greater than 0.8, 1 point for visual prominence greater than 0.7, and the sum of the three scores constitutes the spatial effect score; Add the scores of the above three dimensions and divide by 3 to get the final expressiveness score. For example, "frosted texture" has a brightness analysis score of 3, a clarity analysis score of 2, and a spatial effect score of 1, resulting in an expressiveness score of (3+2+1) / 3=2. "Streamlined shell" has a brightness analysis score of 1, a clarity analysis score of 2, and a spatial effect score of 2, resulting in an expressiveness score of (1+2+2) / 3=1.67.

[0036] Through the scoring of the three dimensions, the system calculates the total score of the importance of each feature. For example: Streamlined shell total score: 3 points (priority) + 2 points (relevance) + 1.67 points (expressiveness) = 6.67 points; Frosted texture total score: 3 points (priority) + 2 points (relevance) + 2 points (expressiveness) = 7 points; Noise reduction performance total score: 1 point (priority) + 1 point (relevance) + 1 point (expressiveness) = 3 points.

[0037] Step 105: Select the features with importance score greater than the score threshold as the product key features.

[0038] Specifically, the score threshold is set to 6 points in the embodiments of the present application. All features in the feature candidate set are screened, and the features with importance scores exceeding the score threshold are determined as product key features. This threshold-based screening mechanism ensures that the finally selected key features not only have sufficient show value, but also effectively highlight the core advantages of the product.

[0039] Step 30: loading a dual-path parallel picture generation system including a first path and a second path, wherein the first path generates a product close-up sequence based on product materials and product key features, and the second path generates a scene atmosphere sequence based on use scenarios and product key features.

[0040] In the embodiments of the present application, the dual-path parallel picture generation system refers to a system architecture that simultaneously runs two picture generation channels. The first path is a product close-up channel, which focuses on the detailed display of the product itself. This channel directly calls product materials (such as product modeling files, product real shooting materials, etc.), and generates a series of close-up sequences that highlight product details based on the selected product key features (such as the streamlined shell of the earphone, the frosted texture, etc.). The second path is a scene atmosphere channel, which focuses on the scene application display of the product. This channel generates a series of picture sequences that reflect the use scenarios and atmosphere of the product based on the pre-set use scenarios (such as indoor office scenarios, home scenarios, etc.) and product key features.

[0041] Based on the above embodiments, as an optional embodiment, the step of generating a product close-up sequence based on product materials and product key features in the first path and generating a scene atmosphere sequence based on the use scenarios and product key features in the second path can include the following steps: Step 201: determining key pictures corresponding to product key features from product materials.

[0042] Specifically, first, the product video in the product material is analyzed and processed. The system uses video frame analysis technology to decompose the product video into a continuous video frame sequence according to a fixed frame rate (such as 24 frames per second). For each video frame, the system uses computer vision algorithms for picture analysis, including identification and extraction of product contour, material details and other information in the picture. The system identifies the product key features as a search condition and preliminarily filters out the video frames containing the features as candidate pictures in the video frame sequence. For example, for the feature of "streamlined shell", the system detects whether the complete contour line of the earphone shell is included in the video frame, and the video frame containing the complete contour is taken as the candidate picture.

[0043] The system scores each candidate picture as follows: The first is the feature area score: calculate the pixel ratio of the feature area in the picture, more than 30% gets 1 point, 15%-30% gets 0.5 point, less than 15% gets 0 point; analyze the definition of the feature area, use Laplacian operator to calculate the definition value, value greater than threshold 100 gets 1 point, 50-100 gets 0.5 point, less than 50 gets 0 point; detect whether the feature area is blocked, no block gets 1 point, block area less than 20% gets 0.5 point, more than 20% gets 0 point; The second is the overall picture score: detecting the brightness uniformity of the picture, calculating the standard deviation of brightness, and scoring 1 if the standard deviation is less than 30, 0.5 if the standard deviation is between 30 and 50, and 0 if the standard deviation is greater than 50; analyzing the picture composition, scoring 1 if the main content is located near the three-line method, 0.5 if the three-line method deviates by less than 10%, and 0 if the three-line method deviates by more than 10%; evaluating the picture stability, scoring 1 if the displacement between adjacent frames is less than 10 pixels, 0.5 if the displacement is between 10 and 20 pixels, and 0 if the displacement is greater than 20 pixels.

[0044] For each candidate picture, the system ranks the sum of the feature region score and the overall picture score in descending order, and selects the candidate picture with a score not lower than the preset score as the key picture. For example: for the "streamlined shell" feature, the 143rd frame has a total score of 5.5 (feature region score 3, overall picture score 2.5), the 127th frame has a total score of 4 (feature region score 2.5, overall picture score 1.5), and the 198th frame has a total score of 4.5 (feature region score 2, overall picture score 2.5). The scores of these three frames are all not lower than the preset score of 4, so they are all key pictures showing the "streamlined shell" feature.

[0045] Step 202: Determine the picture display parameters corresponding to each key picture based on the importance score of the product key feature. The picture display parameters include the picture display time, the picture display ratio, and the picture switching mode.

[0046] Among them, the picture display parameters refer to three core technical parameters for controlling the specific presentation mode of the key picture in the video: picture display time: refers to the duration of each key picture in the video; picture display ratio: refers to the size ratio of the key picture in the video frame; picture switching mode: refers to the transition effect between adjacent key pictures.

[0047] Based on the above embodiment, as an optional embodiment, the step of determining the picture display parameters corresponding to each key picture based on the importance score of the product key feature further includes the following steps: Step 2021: Construct a three-dimensional display parameter table according to the importance score. The horizontal coordinate of the display parameter table is the time coefficient, the vertical coordinate is the ratio coefficient, and the depth coordinate is the switching type.

[0048] Specifically, the system constructs a three-dimensional display parameter table for determining the picture display parameters. The parameter table adopts a three-dimensional coordinate system structure, in which the horizontal coordinate represents the time length coefficient (value range 0.5-2.0), the vertical coordinate represents the proportion coefficient (value range 0.2-1.0), and the depth coordinate represents the switching type (including smooth transition, hard cut, fade-in and fade-out, push, and other types). The system uses a mapping rule to convert the importance score (1-10 points) into a specific parameter value: the time length coefficient is calculated by the formula "0.5+(importance score-1)×0.167", the proportion coefficient is calculated by the formula "0.2+(importance score-1)×0.089", and the switching type is determined according to the score interval (1-2 points correspond to hard cut, 3-4 points correspond to push, 5-6 points correspond to fade-in and fade-out, 7-8 points correspond to smooth transition, and 9-10 points correspond to 3D rotation transition). When there are multiple key features in the product, the system first arranges the features in descending order of importance score to determine the display order, and then uses the parameter decreasing rule: the time length coefficient of adjacent features is reduced by 10% based on the previous level, the proportion coefficient is reduced by 15% based on the previous level, and the switching type is directly mapped according to the respective importance score. For example, for features A, B, and C with importance scores of 8, 6, and 4 points respectively, the parameters of feature A are {time length coefficient 1.67, proportion coefficient 0.82, smooth transition}, the parameters of feature B are {time length coefficient 1.50, proportion coefficient 0.70, fade-in and fade-out}, and the parameters of feature C are {time length coefficient 1.35, proportion coefficient 0.60, push}. In this way, the system not only ensures the accuracy of parameter mapping for a single feature, but also realizes the reasonable decrease of display parameters between multiple features, making the overall display effect more coordinated.

[0049] Step 2022: Based on the visual complexity of each key picture, the time length coefficient is corrected to obtain the picture stay time length of each key picture.

[0050] Specifically, the system analyzes the visual complexity of each key picture, calculates the detail density, color change degree, and content richness, etc. in the picture. The specific calculation method is: divide the picture into multiple grid units, count the number of edge feature points, color change degree, and texture complexity in each unit, and comprehensively obtain the visual complexity index (range 0-1). Then, the system uses the correction formula: actual stay time length=baseline time length×time length coefficient×(1+visual complexity index×0.5) to adjust the original time length coefficient. For example, a product detail picture with a visual complexity index of 0.8, whose original time length of 2 seconds will be corrected to 2.8 seconds, ensuring that the audience has sufficient time to understand the complex picture content.

[0051] Step 2023: Calculate the area ratio of the main body region and the background region in each key picture, and determine the picture display proportion of each key picture combined with the proportion coefficient.

[0052] Specifically, first, the system uses an image segmentation algorithm to identify the subject area and the background area in the key frame, and calculates the pixel area of the two areas. Then, the system calculates the area ratio: ratio = subject area / total frame area. The system combines this ratio with the proportion coefficient in the display parameter table, and uses the formula: final display ratio = baseline ratio x proportion coefficient x (1 + area ratio correction coefficient) to determine the actual display ratio of the frame. For example, for a product close-up frame with a subject area ratio of 0.6, if its proportion coefficient is 0.8, the final display ratio in the video may be set to 75% of the original size, which ensures clear display of the subject content and retains appropriate frame white space.

[0053] Step 2024: Analyze the visual similarity of adjacent key frames, and set the corresponding frame transition mode in combination with the transition type setting.

[0054] Specifically, the system uses image feature matching technology to calculate the visual similarity between adjacent key frames, including color distribution similarity, composition structure similarity, and content theme similarity. The specific calculation uses the feature vector distance method: extract the color histogram, edge feature, and texture feature of the two frames, calculate the Euclidean distance between the feature vectors, and obtain the comprehensive similarity score (range 0-1). The system adjusts the transition type in the display parameter table based on the similarity score: when the similarity score is greater than 0.8, adjust the transition type in the display parameter table to a smoother transition mode (such as hard cut to push, push to fade in and out, fade in and out to smooth transition, smooth transition to 3D rotation transition); when the similarity score is between 0.4 and 0.8, directly use the transition type in the display parameter table; when the similarity score is less than 0.4, adjust the transition type in the display parameter table to a more abrupt transition mode (such as 3D rotation transition to smooth transition, smooth transition to fade in and out, fade in and out to push, push to hard cut). For example, a feature corresponds to a fade-in and fade-out transition type in the display parameter table, if the similarity score of its adjacent frame is 0.9, the final transition mode is adjusted to smooth transition; if the similarity score is 0.3, the final transition mode is adjusted to push. This way of dynamically adjusting the transition type based on visual similarity not only considers the importance of the feature, but also ensures the natural and smooth transition of the frame, improving the visual quality of the video.

[0055] Step 203: Time sequence arrangement of key frames according to frame display parameters, generating product close-up sequence.

[0056] Specifically, the embodiment of the present application designs a set of timing arrangement algorithm based on picture display parameters, which is used to generate a complete product close-up sequence. First, the system constructs an initial time axis according to the picture staying time of each key picture, and arranges all key pictures in the logical sequence of product features. Using a timing optimization algorithm, set the total time target to 30 seconds, set the exact start time and end time for each key picture, if it exceeds, compress the picture time in proportion, if it is insufficient, appropriately extend the staying time of important pictures.

[0057] Next, the system plans the spatial layout based on the picture display ratio, designs the specific presentation position and size of the picture in the video. The system converts each key picture into a video segment of a specific resolution according to its display ratio parameter. For example, a picture with a display ratio of 0.8 will be scaled to 80% of the size of a 1920x1080 video frame. For the transition between adjacent pictures, the system applies a preset picture switching method and inserts the corresponding transition effect at the switching point.

[0058] Finally, the system uses a video synthesis engine to combine the processed key pictures in the order of the time axis to generate a product close-up sequence. The system first performs quality enhancement processing on each key picture, including sharpening, noise reduction, and color correction operations. Then, the system adds a preset transition effect between pictures, such as a 2-second transition transition effect when switching from a product overall appearance picture to a local detail picture. For the product main body in the picture, the system will also add a dynamic tracking special effect to ensure that the product is always in the focus position of the picture.

[0059] In addition, the system will also add auxiliary visual elements to the product close-up sequence, such as product feature annotations, key area local magnification frames, etc., to enhance the display effect of product features. For example, when displaying the touch area of the earphone, the system will add a dynamic indicator line in the picture to highlight the position and method of touch operation. After all the processing is completed, the system will generate a product close-up sequence containing complete transition effects, dynamic effects and auxiliary elements. This parameter-based timing arrangement method not only ensures the coherence and attractiveness of the product close-up sequence, but also enhances the display effect of product features through various visual enhancement methods.

[0060] Step 204: Based on the use scene, the corresponding scene constituent elements are obtained in the preset scene knowledge base, and a basic mood scene sequence is generated based on the scene constituent elements.

[0061] Specifically, the system pre-constructs a structured scene knowledge base containing the constituent elements of various typical use scenarios. The scene constituent element main body includes three levels: environmental elements (such as lighting conditions, spatial size, architectural style), functional elements (such as furniture arrangement, device configuration, use flow), and atmosphere elements (such as color matching, material effect, environmental sound effect). The system retrieves the highest matching scene element combination from the knowledge base according to the product's use scenario type. For example, for a home smart earphone product, the system will select a modern minimalist indoor environmental element (such as a pure color wall, a large area of floor-to-ceiling windows), a home functional element (such as a desk, a sofa, a leisure corner), and a warm and comfortable atmosphere element (such as warm-toned light, soft fabric texture). Then, the system uses a scene generation engine to create a three-dimensional scene model based on the selected scene constituent elements, and generates a serialized display scheme for multiple scene perspectives through a scene arrangement algorithm, forming a basic mood scene sequence. This knowledge base-based scene generation method ensures the authenticity and applicability of the scene.

[0062] Step 205: Determine scene rendering parameters according to product key features, and render the basic mood scene sequence according to the scene rendering parameters to obtain the scene mood sequence.

[0063] Specifically, the system analyzes the visual attributes of the product's key features, including material properties (such as metal luster, frosted texture), form characteristics (such as curved modeling, geometric structure), and functional characteristics (such as detachability, interaction mode). Based on these feature attributes, the system sets corresponding rendering parameters, including lighting parameters (such as main light source position, lighting intensity, shadow softness), material parameters (such as reflectivity, roughness, transparency), and environmental parameters (such as depth of field effect, color saturation, atmospheric effect). For example, to highlight the metal wire drawing process of the earphone, the system will set a directional point light source and increase the anisotropic reflectivity parameters of the material; to demonstrate the ergonomics design of the product, the system will increase the environmental light shading effect and enhance the stereoscopic sense of the product outline. Then, the system uses a rendering engine to apply these rendering parameters to each scene picture of the basic mood scene sequence, achieving high-quality visual effects through rendering. Finally, the system performs post-color processing on the rendered scene sequence to unify the picture style, forming the final scene mood sequence. This feature-oriented rendering processing method ensures the perfect integration of product features and scene atmosphere, enhancing the artistic expressiveness of the video.

[0064] Step 40: Perform picture segmentation on the product close-up sequence and the scene mood sequence to generate a visual element library.

[0065] In the embodiments of the present application, the visual element library refers to a paired visual material set composed of product display segments and corresponding scene display segments. Each paired combination is referred to as a visual element pair, and is accompanied by corresponding annotation information. It can be further understood that when a certain feature or function of a product is displayed, a matching scene environment needs to be selected to enhance the display effect. The combination of product display content and scene display content constitutes a visual element pair.

[0066] In a feasible embodiment, the product close-up sequence and the scene mood sequence are analyzed in the time axis to determine the time nodes of key frames. Then, the system uses a frame segmentation algorithm to segment the two sequences at these time nodes to obtain multiple independent display segments. For the product close-up sequence, the system extracts product display segments containing single feature display; for the scene mood sequence, the system extracts corresponding scene display segments. Then, the system pairs each product display segment with its time corresponding scene display segment to form a visual element pair. Finally, the system adds annotation information to each visual element pair, including scene type, product feature, matching correlation degree and other attributes, and stores all annotated visual element pairs in the visual element library; this segmentation and pairing method ensures accurate correspondence and efficient reuse of visual elements.

[0067] On the basis of the above-mentioned embodiments, as another optional embodiment, the step of segmenting the product close-up sequence and the scene mood sequence to generate the visual element library can include the following steps: Step 301: identifying frame pause points and transition points in the product close-up sequence.

[0068] Specifically, the system analyzes the product close-up sequence frame by frame, and identifies frame change characteristics by calculating the image difference value between adjacent frames. The frame difference method is used to calculate the pixel difference degree of each pair of adjacent frames. For example, when the difference degree is less than a preset threshold (such as 0.05) and the duration exceeds a preset time length (such as 0.5 seconds), the position is marked as a frame pause point. At the same time, the system detects the frame transition effect characteristics (such as the characteristic parameters of fade-in and fade-out, push, dissolve and other effects), and identifies the specific positions of each transition point. For example, when typical transition characteristics such as brightness gradient or frame transition are detected, the system will mark the time point as a transition point. This accurate point identification method provides a reliable reference point for subsequent sequence segmentation.

[0069] Step 302: segmenting the product close-up sequence at each frame pause point and each transition point to obtain multiple product display segments.

[0070] Specifically, the system arranges all the identified points in chronological order to form a segmented timeline. For each segment, the system determines its start and end time points and performs an accurate cut on the product close-up sequence at these time points. The system uses lossless segmentation techniques to ensure that there is no picture damage or frame loss at the segmentation points. For example, when segmenting at a certain picture pause point (time point 15.5 seconds), the system accurately locates the nearest key frame and completes the segmentation operation at that frame, generating an independent product display segment. This key point-based segmentation method ensures the integrity and continuity of each product display segment.

[0071] Step 303: According to the time interval of the product display segment, the corresponding scene display segment is extracted from the scene atmosphere sequence.

[0072] Specifically, the time interval information (start and end time points) of each product display segment is mapped to the timeline of the scene atmosphere sequence. Then, the system accurately extracts the scene atmosphere sequence within the corresponding time interval. To ensure accuracy, the system uses a frame-level synchronization mechanism, i.e., accurately matching the time code of each frame during extraction. For example, when the time interval of a product display segment is 15.5 seconds to 18.2 seconds, the system accurately locates this time interval in the scene atmosphere sequence and ensures that the extracted scene display segment completely matches the product display segment in terms of duration. This accurate time-synchronous extraction method ensures perfect coordination between product display and scene display.

[0073] It should be noted that the scene atmosphere sequence and the product close-up sequence have a one-to-one correspondence. In the previous step, the system has generated a matching scene atmosphere sequence based on the product key features, i.e., each product display picture has its corresponding scene supporting picture. Therefore, when the system extracts the scene display segment from the scene atmosphere sequence according to the time interval of the product display segment (e.g., 15.5 seconds to 18.2 seconds), the extracted scene corresponds to the product display content. For example, when the product display segment displays earphone touch operation, the scene display segment within the corresponding time interval displays the user using the touch function in the corresponding scene, ensuring the content coordination between product display and scene display.

[0074] Step 304: Each product display segment and the corresponding scene display segment form a visual element pair, and each visual element pair is labeled to generate a visual element library.

[0075] Specifically, each product display segment is paired with its corresponding scene display segment to form a complete visual element pair. Then, the system performs multi-dimensional labeling on each visual element pair, including content labeling (such as displayed product features, scene types), technical labeling (such as picture resolution, color parameters), and correlation labeling (such as the matching degree of the product and the scene). For example, for a visual element pair displaying the touch control function of earphones, the system will label information such as "product type: smart earphones; display feature: touch operation; scene type: home office; picture parameters: 1080p / 30fps; matching degree: high". Finally, the system organizes all the labeled visual element pairs according to the preset classification system and stores them in the visual element library; this structured organization method makes the visual element library have good retrievability and reusability.

[0076] Step 50: Generating a narrative script and an audio script according to the group characteristics of the target audience group and the key features of the product, and arranging the visual element library according to the narrative script to generate a picture display sequence.

[0077] In the embodiments of the present application, the narrative script refers to a standardized product introduction script generated according to the key features of the product, which contains the description content, display order, and importance of each product key feature.

[0078] The audio script refers to a dubbing script and sound effect design scheme converted from the narrative script, including commentary text for voice synthesis and configuration information of various sound effect elements. For example, based on the description content "supports active noise reduction function" in the narrative script, the audio script will give the specific commentary "equipped with the new generation of active noise reduction technology, which can effectively reduce 95% of environmental noise", and specify the background sound effect type that needs to be matched during this segment of dubbing.

[0079] The picture display sequence refers to a video sequence containing product close-up pictures and supporting scene pictures generated according to the audio script, where the display content, duration, and transition method of each picture are synchronized with the audio script. For example, when the audio script plays the noise reduction function commentary, the picture display sequence will correspondingly display the demonstration picture and the actual use scene picture of the noise reduction function.

[0080] Based on the above embodiments, as an optional embodiment, the step of generating a narrative script and an audio script according to the group characteristics of the target audience group and the key features of the product can include the following steps: Step 401: Locate each product key feature in the product material to obtain the corresponding promotional text, and each visual element pair in the visual element library includes at least one promotional text.

[0081] Specifically, the system performs text analysis on product materials such as product manuals, promotional brochures, official websites, etc. using natural language processing techniques to identify descriptions related to product key features. The system uses keyword matching and semantic relevance analysis methods to map the identified text content to pre-defined product key features. For example, when the text "adopting a dual-drive unit design" is identified, the system will locate it as promotional text for acoustic features. Then, the system associates the located promotional text with visual element pairs in the visual element library to ensure that each visual element pair contains at least one corresponding promotional text. This feature positioning method ensures accurate communication of product information.

[0082] Step 403: Determine group characteristics based on historical behavior data of target audience groups.

[0083] Specifically, the system collects two types of historical behavior data: target audience group behavior data on similar target marketing products (such as purchase records, usage evaluations, and after-sales feedback of different brands of Bluetooth earphones) and group behavior data on other products of the same brand (such as website browsing records, social media interactions, and promotion activity participation). Then, from these comprehensive behavior data, user characteristics are extracted, including demographic characteristics (such as age distribution, occupation composition), consumer behavior characteristics (such as price sensitivity, brand loyalty), usage habit characteristics (such as scene preference, function appeal), etc. Based on these data, the system generates multi-dimensional group characteristic labels to provide accurate guidance for subsequent content customization presentation. For example, through analysis, it is found that the target audience group pays more attention to noise reduction performance and battery life on similar earphone products, and prefers simple and intuitive data display methods. The system will adjust the display strategy of related content accordingly. This comprehensive historical data-based group profiling ensures that marketing content can more accurately reach the core needs of the target audience.

[0084] Step 404: Determine the display parameters of promotional text in each visual element pair according to group characteristics.

[0085] In one possible embodiment, the system first establishes a mapping relationship library between group characteristics and display parameters, including font selection, font size, display duration, animation effects, etc. For example, for young groups (25-35 years old), the system will select modern and simple style fonts with dynamic display effects; for professional user groups, the system will increase the display duration of technical parameters and use professional terms. Then, the system sets appropriate display parameters for promotional text in each visual element pair according to group characteristics. For example, when displaying earphone noise reduction performance, for business people groups, the system will set "noise reduction effect data" to be presented in the form of prominent numbers with a gradual entry animation effect. This parameterized display design improves the effectiveness of information communication.

[0086] In another possible embodiment, a machine learning method can be employed to dynamically generate display parameters. The system collects a large number of data samples of historical marketing videos, including display effect data (such as completion rate, interaction rate, conversion rate, etc.) of various types of promotional text under different group characteristics. Then, the system uses deep learning algorithms to train a display parameter prediction model. The model takes group characteristics (such as age, occupation, and consumption ability) as input to predict the optimal combination of display parameters. For example, when identifying the target group as "25-35 years old, high education, and focusing on professional performance", the model will automatically recommend the most suitable font style, display duration, and animation effect for this group based on historical data. For promotional text of earphone noise reduction performance, the model may suggest using professional terminology, with data visualization effects, and giving a longer display duration. The system also collects real-time parameter effect feedback to continuously optimize the accuracy of the model. This machine learning-based dynamic parameter generation method improves the intelligence and accuracy of display parameter configuration.

[0087] Step 405: generate a corresponding audio script based on the narration script.

[0088] In a possible embodiment, the narration script is subjected to semantic analysis to identify key content and tone features. Then, the system uses speech synthesis algorithms to convert the text content into suitable voice scripts for broadcasting, including speech speed, tone, emphasis, and other details. For example, when the narration script describes "using the new generation of active noise reduction technology", the system will set a slower speech speed and emphasis tone in the audio script. At the same time, the system will also configure the corresponding background music and sound effects according to the content characteristics. For example, when demonstrating the noise reduction function, configure a gradual transition from noisy to quiet sound effects to highlight the noise reduction effect. This multi-dimensional audio design enhances the appeal of product demonstration.

[0089] On the basis of the above-mentioned embodiments, as another optional embodiment, the step of generating a corresponding audio script based on the narration script can further include the following steps: Step 4051: determine the main audio and secondary audio of each visual element pair in the narration script, wherein the main audio at least includes voice dubbing of promotional text or product sound effects, and the secondary audio at least includes background music or environmental sound effects.

[0090] Specifically, the system uses an audio level division algorithm to analyze the audio requirements of each visual element pair in the narrative script, dividing it into two levels: main audio and secondary audio. For the determination of main audio, the system analyzes the content type and importance of the promotional text: when the promotional text describes the core function of the product (such as "48-hour battery life"), the voice dubbing is set as the main audio; when the product operation effect (such as touch operation) is shown, the corresponding operation sound effect is set as the main audio. For specific product functions (such as noise-cancelling earphones), the system also generates special product sound effects, such as the comparison of noise-cancelling before and after. In specific implementation, the system uses a semantic analysis algorithm to identify keywords and key sentence patterns in the promotional text, and assigns audio priority accordingly. For example, content containing keywords such as "patented technology" and "innovation breakthrough" will be prioritized as main audio.

[0091] For secondary audio, the system plans based on the scene attributes and emotional characteristics of the visual element pair. First, the system extracts environmental features such as "indoor office" and "outdoor sports" from the scene annotation information, and matches the corresponding environmental sound effects. Then, it analyzes the emotional tone of the scene and selects the appropriate background music style. For example, when showing a business office scene, the system will choose low-key environmental sound effects (such as keyboard typing and office environment) as secondary audio, and match them with soothing light music as a backdrop. This hierarchical audio planning ensures that the sound elements are highlighted and layered, improving the professionalism of the overall audio effect.

[0092] Step 4052: determining the first voice dubbing parameter of the main audio from the preset voice dubbing library according to the group characteristics, and determining the second voice dubbing parameter of the secondary audio from the preset voice dubbing library according to the use scenario.

[0093] Specifically, the system uses a two-dimensional voice dubbing parameter matching algorithm to analyze the group characteristics in multiple dimensions, including basic attributes (age, occupation, etc.), consumer characteristics (purchasing power, brand preference), and usage habits (scene preference, function appeal), etc. Based on these characteristics, the system matches the first voice dubbing parameter from the preset voice dubbing library. For example, when the target group is "25-35 years old, business people, and focuses on professional performance", the system will choose a voice dubbing parameter combination with a speed of 160-180 words per minute, a low-mid tone line, and a professional expression tone. Among them, the best voice dubbing parameter combination can be determined by feature weight calculation: speed weight 0.3, tone weight 0.4, tone weight 0.3, and according to the scores of each dimension characteristic, the final voice dubbing parameter combination with the highest score is calculated.

[0094] Meanwhile, the system analyzes the use scene features to determine the second dubbing parameters from the preset dubbing library. For example, when the "office use scene" is recognized, the system will select soft background music (volume is 30% of the main audio), configure moderate indoor reverb effect (reverb time 1.0 seconds), and add slight office environment sound effects (such as keyboard sound, air conditioner sound, etc.). The system intelligently matches the scene attributes with the audio parameters through the scene feature analysis engine to ensure that the secondary audio effect can accurately highlight the product use environment. For example, when demonstrating the office scene of the noise reduction earphone, the system will first set a strong environmental noise, then demonstrate the noise reduction effect through sound fading effect, and cooperate with comfortable background music to highlight the product performance.

[0095] This two-dimensional dubbing parameter matching method not only ensures the accurate transmission of the main audio, but also enhances the scene immersion through reasonable secondary audio configuration, thereby improving the professionalism and appeal of the overall audio effect.

[0096] Step 4053: Determine the mixing parameters of each visual element pair according to the annotations of each visual element pair, combined with the first dubbing parameters and the second dubbing parameters.

[0097] Specifically, the system analyzes the annotation information of the visual element pair, including content importance (1-10 points), scene type (such as indoor, outdoor), emotional tone (such as professional, relaxed), etc. Then, the system integrates these annotation information with the first dubbing parameters and the second dubbing parameters to generate mixing parameters. For example, when the visual element pair demonstrates the core function of the noise reduction earphone, the annotation shows "importance 9 points, indoor office scene, professional experience", the system will set the mixing parameters as follows: the main audio (professional male voice dubbing) volume is set to 0.7, the product sound effect (noise reduction effect sound) volume is set to 0.6, the secondary audio (background music) volume is set to 0.3, and the environmental sound effect (office noise) volume is gradually changed from 0.5 to 0.1. The system also sets the sound channel allocation parameters, such as allocating the dubbing mainly in the center channel and dispersing the environmental sound effect in the left and right channels to create an immersive listening experience.

[0098] For audio transition effects, the system configures corresponding fading parameters based on the duration and switching mode of the visual element pair. For example, when the scene transitions from a noisy office to a noise reduction state, the system sets a 1.5-second gradual weakening transition for the environmental sound, and simultaneously cooperates with a 0.8-second fade-in effect for the background music. To ensure the audio hierarchy, the system uses a dynamic priority adjustment mechanism to automatically increase the volume proportion of the main audio during product function demonstration, and to strengthen the presence of the secondary audio during scene transition. This precise mixing parameter configuration ensures the harmony and unity of each audio element, improving the professionalism and appeal of the overall audio effect. At the same time, the system will ensure that there is no frequency conflict between audio tracks through audio spectrum analysis to ensure the clarity of the final mixing effect.

[0099] Step 4054: Integrate each visual element pair and corresponding mixing parameters into an audio script.

[0100] Specifically, each visual element pair and its corresponding mixing parameters are arranged in chronological order to form an initial audio timeline. Then, the system adds audio transition designs between adjacent visual element pairs, including sound fade-in and fade-out, sound effect superposition, etc. The system also checks the overall audio rhythm to ensure that the switching rhythm of the main and secondary audio is consistent with the picture display rhythm. Finally, the system generates a complete audio script document, which records the audio element combination and mixing parameter configuration at each time point in detail. For example, in the noise reduction function display paragraph from 15.5 seconds to 18.2 seconds, the audio script will mark specific configuration information such as "main audio: male voice dubbing (volume 0.7) + noise reduction effect sound (volume 0.6), secondary audio: light music background (volume 0.3)". This structured audio script integration method ensures the standardization and operability of subsequent production.

[0101] Step 60: Generate audio sequences corresponding to picture display sequences according to the audio script.

[0102] Specifically, the embodiments of the present application provide an audio sequence generation engine for converting the audio script into an audio sequence that is precisely synchronized with the picture display sequence. First, the system reads the mixing parameters of each visual element pair from the audio script, including volume ratio, channel allocation, and transition effect parameters. Then, the system generates audio time codes according to the mixing parameters to determine the precise playback time points, duration, and transition methods of each audio element. For example, when the audio script shows "main audio: male voice dubbing (volume 0.7) + noise reduction effect sound (volume from 0.5 linearly fades to 0.1), secondary audio: light music background (volume 0.3)", the system will generate the corresponding time codes: "T15.5-18.2: main audio dubbing executes 0.7 times volume; T15.5-17.0: noise reduction effect from 0.5 linearly fades to 0.1; T15.5-18.2: background music maintains 0.3 times volume".

[0103] Then, the system synthesizes each audio element according to the mixing parameters based on the generated time codes to generate a complete audio sequence. Audio processing algorithms can be used to perform mixing operations on audio tracks according to time code nodes, including volume adjustment, channel allocation, and sound effect superposition. For example, in the noise reduction function display segment, the system strictly executes the gradual weakening effect of environmental noise according to the time code, while maintaining the stable volume of the dubbing and background music, ensuring the precise synchronization of the audio sequence with the picture display sequence. This time code-based audio sequence generation method ensures the accuracy and professionalism of audio processing, and improves the auditory experience of product display.

[0104] Step 70: Time sequence combination of the picture display sequence and the audio sequence, generation of the marketing video of the target marketing product, and display of the marketing video.

[0105] Specifically, first, the system aligns the time markers in the picture display sequence and the audio sequence through time code matching technology. The system reads the time code information of the two sequences, establishes a time mapping relationship table, and ensures the synchronization of audio and video at each key node. For example, when the audio sequence appears at T15.5 seconds, the system will accurately correspond it to the noise reduction function display picture at T15.5 seconds in the picture display sequence.

[0106] Then, the system uses a professional video encoder to synthesize the aligned picture display sequence and audio sequence to generate a marketing video. H.264 video encoding and AAC audio encoding standards can be used, with a video code rate of 8 Mbps and an audio sampling rate of 48 kHz to ensure the output quality of 1080p high-definition video. During the encoding process, the system uses key frame technology to ensure the smoothness of picture switching and audio-video synchronization buffering mechanism to prevent audio-visual synchronization phenomenon. For example, when displaying the noise reduction function of earphones, the fade-in of the noise reduction effect sound must be accurately corresponding to the visual change of the noise reduction state.

[0107] Finally, the system displays the generated marketing video through multi-platform adaptation technology. The system will automatically adjust the resolution, code rate and playback format of the video according to the requirements of different playback platforms (such as mobile, PC, large screen display, etc.). At the same time, the system will also monitor the playback effect and collect user viewing data, including completion rate, interaction rate and other indicators, for subsequent optimization. This precise time sequence combination and multi-platform display method ensures that the marketing video can present professional audio-visual effects in various scenarios, effectively conveying the product value.

[0108] It should be noted that in another possible embodiment, a main marketing video can be generated for display, and multiple backup marketing videos can be generated for the customer to choose from. The system first generates a main marketing video that highlights the core selling points of the product based on the key features of the product and the main portrait features of the target audience group. For example, for a smart earphone product, the main marketing video highlights the core functional features such as "noise reduction, long battery life, AI voice". At the same time, the system generates 3-5 different backup marketing videos based on different audience group preferences and use scenarios. For example, for the business group, the system generates a marketing video that highlights the noise reduction effect and call quality (video 1); for the sports group, the system generates a marketing video that emphasizes the waterproof performance and wearing stability (video 2); for the music group, the system generates a marketing video that highlights the sound quality and sound effect (video 3). These backup videos show the same product, but each has its own focus on content, presentation and style.

[0109] See Figure 3 An exemplary marketing video display interface is provided for the embodiments of the present application. The interface mainly displays the marketing video content of the smart earphone product, in which the main visual area shows the product form of the split earphone and the modeling features of the charging storage box through simple line drawings, and marks the key feature words of the product: "noise reduction, long battery life, AI voice". The video playback control module is set in the middle of the interface, including a progress bar and a play button, with an interactive prompt "click to play video". Below the interface, three playback entrances of backup marketing videos are set under the title "smart earphone", marked as "video 1", "video 2" and "video 3", and the user can click any one of the backup videos to watch and choose the final marketing video. The user comment interaction area is configured at the bottom. The interface effectively displays the marketing video content and provides a convenient user interaction method through a clear hierarchical structure and simple visual language.

[0110] On the basis of the above embodiments, as an optional embodiment, if the user is not satisfied with the generated marketing video, feedback and suggestions can be made, and then the system generates a new marketing video based on the further input content or suggestions of the user.

[0111] See Figure 4Another exemplary user interaction interface diagram provided for the embodiments of the present application. This interface demonstrates the interaction process when the user provides feedback on the marketing video. When the user is not satisfied with the generated marketing video, the system will guide the user to provide specific feedback through a dialog box, such as "Does the current marketing video meet your needs? Welcome to provide your suggestions" displayed at the top of the interface. The user can directly provide optimization suggestions or provide feedback suggestions based on viewing the other three backup videos, such as the user's feedback suggestion: "Based on Video 2, highlight the appearance of the earphone and its color value, and update the marketing video", the system further generates a status prompt of "Generating, please wait" with a loading animation indicator. The search input box is set at the bottom of the interface to facilitate the user to input more feedback at any time.

[0112] This interaction design reflects the intelligent iteration capability of the system, which collects the specific needs and suggestions of the user and adjusts and optimizes the content of the marketing video accordingly, so as to generate marketing materials that meet the user's expectations. The system will adjust and optimize the video content based on the user's feedback, such as highlighting specific product features or use scenarios, to provide marketing video content that better meets the user's needs.

[0113] Based on the above embodiments, as an optional embodiment, the system provides diversified video preview and interaction functions, not only supporting basic playback control of the video, but also constantly optimizing the generated marketing video based on user feedback, such as the system generating a new marketing video based on user feedback.

[0114] Specifically, first, extract the keywords in the user feedback information, and dynamically update the marketing video content and script based on these keywords. For example, when the system receives user feedback such as "hope to highlight the small and portable features of the earphone, suitable for carrying around", first extract the keywords "small", "portable", "carry around" through semantic analysis. The system then matches the corresponding visual elements from the material library according to the preset keyword mapping rules. For example, for the keywords "small" and "portable", the system will select to display the comparison picture of the earphone and the size of the palm, or display the scene of the earphone easily put into the pocket, directly showing the portability of the product.

[0115] In terms of advertising script generation, the system uses a template-based script combination method. For example, when "small", "portable" and other keywords are extracted, the system will select the appropriate expression from the preset script template library, such as "small and exquisite, light and quality". If "fashion", "appearance" and other keywords appear in the user feedback, the system will generate scripts such as "color value online, very cool to wear" to highlight the product's appearance design.

[0116] Meanwhile, the system reorders the video content according to the weights of the keywords. For example, when the user emphasizes portability, the system adjusts the pictures showing the portability of the product to the front end of the video and adjusts the picture duration accordingly to highlight the key features that the user is concerned about. In this way, the system can intelligently adjust the marketing content according to user feedback and provide a display effect that better meets the user's needs.

[0117] See Figure 5 Another exemplary marketing video display interface provided by the embodiment of the present application.

[0118] The interface uses a simple line diagram to present product features through two groups of key visual elements; For example, in terms of product appearance display, the left side shows a line diagram of a separate earphone, and the right side shows a schematic diagram of a charging storage box, which outlines the overall outline of the product through simple lines. The top text is "small and delicate, light and quality", and the bottom text is "color value is online, and it is very cool to wear". The text highlights the keywords "small", "delicate", "light", "quality", and "color value".

[0119] See Figure 6 A module schematic diagram of a marketing video generation system provided by the embodiment of the present application, wherein the system comprises: A production instruction acquisition module for acquiring production instructions of a target marketing product, the production instructions at least including product materials, use scenarios, target audience groups, and additional requirements; A key feature determination module for determining product key features according to the additional requirements and the product materials; A parallel sequence generation module for loading a double-path parallel picture generation system including a first path and a second path, wherein the first path generates a product close-up sequence based on the product materials and the product key features, and the second path generates a scene mood sequence based on the use scenarios and the product key features; An element library generation module for segmenting the product close-up sequence and the scene mood sequence to generate a visual element library; A script generation module for generating a narrative script and an audio script according to the group characteristics of the target audience groups and the product key features, and arranging the visual element library according to the narrative script to generate a picture display sequence; An audio sequence generation module for generating an audio sequence corresponding to the picture display sequence according to the audio script; The marketing video generation module is configured to combine the picture display sequence and the audio sequence in time sequence, generate a marketing video of the target marketing product, and display the marketing video.

[0120] Optionally, the key feature determination module is further configured to extract product basic information from the product material, the product basic information including function description information in product introduction text, appearance feature information in product images, and use effect information in product videos. The product basic information is classified according to function dimensions, appearance dimensions, and use effect dimensions to generate a feature candidate set. The additional demand is subjected to semantic analysis to identify a user key display intention. The features in the feature candidate set are subjected to importance scoring based on the user key display intention, the importance scoring including display priority score, feature correlation degree score, and expressiveness score. Features with importance scores greater than a score threshold are selected as product key features.

[0121] Optionally, the parallel sequence generation module is further configured to determine key pictures corresponding to the product key features from the product material. Picture display parameters corresponding to each of the key pictures are determined based on the importance scores of the product key features, the picture display parameters including picture staying time length, picture display proportion, and picture switching mode. The key pictures are time-sequentially arranged according to the picture display parameters to generate a product close-up sequence. Corresponding scene constituting elements are obtained in a preset scene knowledge base based on the use scenario, and a basic mood scene sequence is generated based on the scene constituting elements. Scene rendering parameters are determined according to the product key features, and the basic mood scene sequence is rendered according to the scene rendering parameters to obtain a scene mood sequence.

[0122] Optionally, the parallel sequence generation module is further configured to construct a three-dimensional display parameter table according to the importance scores, the horizontal coordinate of the display parameter table being a time length coefficient, the vertical coordinate being a proportion coefficient, and the depth coordinate being a switching type. The time length coefficient is corrected based on the visual complexity of each of the key pictures to obtain the picture staying time length of each of the key pictures. The area ratio of a subject region to a background region in each of the key pictures is calculated, and the picture display proportion of each of the key pictures is determined in combination with the proportion coefficient. The visual similarity of adjacent key pictures is analyzed, and the corresponding picture switching mode is set in combination with the switching type setting.

[0123] Optionally, the element library generation module is further configured to identify picture pause points and transition points in the product close-up sequence. The product close-up sequence is segmented at the picture pause points and the transition points to obtain a plurality of product display segments. According to time intervals of the product display segments, corresponding scene display segments are intercepted from the scene mood sequence. Each product display segment and a corresponding scene display segment are combined to form a visual element pair, and the visual element pair is labeled to generate a visual element library.

[0124] Optionally, the script generation module is further configured to locate the product key features in the product material to obtain promotional text corresponding to each key feature, and each visual element pair in the visual element library includes at least one promotional text. According to historical behavior data of the target audience group, group features are determined. According to the group features, display parameters of the promotional text in each visual element pair are determined. The visual element pairs and the display parameters of the corresponding promotional text are integrated into a narrative script. Based on the narrative script, a corresponding audio script is generated.

[0125] Optionally, the script generation module is further configured to determine main audio and secondary audio of each visual element pair in the narrative script, wherein the main audio includes at least voice dubbing of the promotional text or product sound effects, and the secondary audio includes at least background music or environmental sound effects. According to the group features, first dubbing parameters of the main audio are determined from a preset dubbing library, and according to the use scenario, second dubbing parameters of the secondary audio are determined from the preset dubbing library. According to the labels of each visual element pair, and in combination with the first dubbing parameters and the second dubbing parameters, mixing parameters of each visual element pair are determined. The visual element pairs and the corresponding mixing parameters are integrated into an audio script.

[0126] It should be noted that: the system provided in the above embodiments, when realizing its functions, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0127] The embodiment of the present application further provides a computer storage medium, which can store a plurality of instructions, the instructions being suitable for being loaded by a processor and performing the marketing video generation method of the above embodiment, and the specific execution process can be referred to the specific description of the above embodiment, which will not be repeated here.

[0128] Please refer to Figure 7 The present application further discloses an electronic device. Figure 7 is a structural schematic diagram of an electronic device disclosed by the embodiment of the present application. The electronic device 300 can include at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0129] The communication bus 302 is configured to realize the connection and communication between the components.

[0130] The user interface 303 can include a display screen (Display) and a camera (Camera), and the optional user interface 303 can further include a standard wired interface and a wireless interface.

[0131] The network interface 304 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0132] The processor 301 can include one or more processing cores. The processor 301 connects various parts of the server through various interfaces and lines, executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Optionally, the processor 301 can be realized in at least one of the hardware forms of a digital signal processing (Digital Signal Processing, DSP), a field programmable gate array (Field-Programmable Gate Array, FPGA), and a programmable logic array (Programmable Logic Array, PLA). The processor 301 can be integrated with a combination of one or more of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU), and a modem. Among them, the CPU is mainly used to process the operating system, user interface and application programs; the GPU is used to render and draw the content to be displayed on the display screen; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 301, but can be realized by a separate chip.

[0133] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 7 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for generating marketing videos.

[0134] exist Figure 7 In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 301 can be used to call an application program stored in the memory 305 for a marketing video generation method. When executed by one or more processors 301, the electronic device 300 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0135] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0136] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the division of the apparatus embodiments is merely illustrative, and the division of units can be changed according to actual conditions, such as a combination or integration of some units, or a deletion of some features, or an addition of some features. In addition, the coupling or direct coupling or communication connection between the shown or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0137] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0138] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0139] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: U disk, mobile hard disk, magnetic disk or optical disk, and various media that can store program codes.

[0140] The above are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the true principles of the present disclosure.

[0141] The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field not described in the present disclosure. The specification and examples are only considered as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for marketing video generation, the method comprising: The method comprises: obtaining production instructions of a target marketing product, the production instructions comprising at least product materials, use scenarios, target audience groups and additional requirements; determining product key features according to the additional requirements and the product materials; loading a double-path parallel picture generation system comprising a first path and a second path, wherein the first path generates product close-up sequences based on the product materials and the product key features, and the second path generates scene mood sequences based on the use scenarios and the product key features; performing picture segmentation on the product close-up sequences and the scene mood sequences to generate a visual element library; generating a narrative script and an audio script according to the group characteristics of the target audience groups and the product key features, and arranging the visual element library according to the narrative script to generate a picture display sequence; generating an audio sequence corresponding to the picture display sequence according to the audio script; performing time sequence combination on the picture display sequence and the audio sequence to generate a marketing video of the target marketing product, and displaying the marketing video.

2. The marketing video generation method of claim 1, wherein, The determination of the product key features according to the additional requirements and the product materials comprises: extracting product basic information from the product materials, the product basic information comprising function description information in product introduction text, appearance feature information in product images and use effect information in product videos; classifying the product basic information according to function dimensions, appearance dimensions and use effect dimensions to generate a feature candidate set; performing semantic analysis on the additional requirements to identify user key display intentions; performing importance scoring on features in the feature candidate set based on the user key display intentions, the importance scoring comprising display priority scores, feature correlation degree scores and expressiveness scores; selecting features with importance scores greater than a score threshold as product key features.

3. The marketing video generation method of claim 2, wherein, The first path generates product close-up sequences based on the product materials and the product key features, and the second path generates scene mood sequences based on the use scenarios and the product key features, which comprises: determining key pictures corresponding to the product key features from the product materials; determining picture display parameters corresponding to each of the key pictures based on the importance scores of the product key features, the picture display parameters comprising picture staying time lengths, picture display proportions and picture switching modes; performing time sequence arrangement on the key pictures according to the picture display parameters to generate product close-up sequences; obtaining corresponding scene constituting elements in a preset scene knowledge base based on the use scenarios, and generating a basic mood scene sequence based on the scene constituting elements; determining scene rendering parameters according to the product key features, and rendering the basic mood scene sequence according to the scene rendering parameters to obtain a scene mood sequence.

4. The marketing video generation method of claim 3, wherein, The determination of the picture display parameters corresponding to each of the key pictures based on the importance scores of the product key features comprises: constructing a three-dimensional display parameter table according to the importance scores, wherein the horizontal coordinates of the display parameter table are time length coefficients, the vertical coordinates are proportion coefficients, and the depth coordinates are switching types. The time length coefficient is corrected based on the visual complexity of each key picture to obtain a picture stay time length of each key picture; An area ratio of a subject area to a background area in each key picture is calculated, and a picture display ratio of each key picture is determined in combination with the ratio coefficient; Visual similarity of adjacent key pictures is analyzed, and a corresponding picture switching mode is determined in combination with the switching type setting.

5. The marketing video generation method of claim 1, wherein, The picture segmentation of the product close-up sequence and the scene mood sequence generates a visual element library, including: Picture pause points and transition points in the product close-up sequence are identified; The product close-up sequence is segmented at each picture pause point and each transition point to obtain a plurality of product display segments; According to the time interval of the product display segment, a corresponding scene display segment is intercepted from the scene mood sequence; Each product display segment and the corresponding scene display segment form a visual element pair, and the visual element library is generated after each visual element pair is labeled.

6. The marketing video generation method of claim 1, wherein, The narrative script and the audio script are generated according to the group characteristics of the target audience group and the product key characteristics, including: Each product key characteristic is located in the product material to obtain a corresponding promotional text, and each visual element pair in the visual element library includes at least one promotional text; Group characteristics are determined according to historical behavior data of the target audience group; Display parameters of the promotional text in each visual element pair are determined according to the group characteristics; Each visual element pair and the corresponding display parameters of the promotional text are integrated into a narrative script; The corresponding audio script is generated based on the narrative script.

7. The marketing video generation method of claim 6, wherein, The corresponding audio script is generated based on the narrative script, including: The main audio and the secondary audio of each visual element pair in the narrative script are determined, wherein the main audio at least includes voice dubbing of the promotional text or product sound effects, and the secondary audio at least includes background music or environmental sound effects; The first dubbing parameter of the main audio is determined from a preset dubbing library according to the group characteristics, and the second dubbing parameter of the secondary audio is determined from a preset dubbing library according to the use scenario; The mixing parameter of each visual element pair is determined according to the label of each visual element pair, in combination with the first dubbing parameter and the second dubbing parameter; Each visual element pair and the corresponding mixing parameter are integrated into an audio script.

8. A marketing video generation system characterized by, The system includes: A production instruction acquisition module is configured to acquire production instructions of a target marketing product, the production instructions including at least product materials, use scenarios, target audience groups, and additional requirements; A key characteristic determination module is configured to determine product key characteristics according to the additional requirements and the product materials; A parallel sequence generation module is configured to load a dual-channel parallel picture generation system including a first channel and a second channel, wherein the first channel generates a product close-up sequence based on the product materials and the product key characteristics, and the second channel generates a scene mood sequence based on the use scenarios and the product key characteristics; An element library generation module is configured to perform picture segmentation on the product close-up sequence and the scene mood sequence to generate a visual element library; A script generation module is configured to generate a narration script and an audio script according to the group characteristics of the target audience group and the product key characteristics, and arrange the visual element library according to the narration script to generate a picture display sequence; An audio sequence generation module is configured to generate an audio sequence corresponding to the picture display sequence according to the audio script; A marketing video generation module is configured to combine the picture display sequence and the audio sequence in time sequence to generate a marketing video of the target marketing product, and display the marketing video.

9. A computer-readable storage medium, characterized in that, A computer readable storage medium stores a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the method of any one of claims 1-7.

10. An electronic device, comprising: An electronic device includes a processor, a memory, a user interface, and a network interface, the memory is configured to store instructions, the user interface and the network interface are configured to communicate with other devices, and the processor is configured to execute the instructions stored in the memory to cause the electronic device to perform the method of any one of claims 1-7.

Citation Information

Cited By

  • Data video generation method and device, equipment and medium

    CN121482220A