Video generation method and system, video medium, medium and program product
By automating the selection of product metadata and materials from e-commerce platforms through a video middleware platform, and combining it with a thought chain reasoning mechanism, the video quality and efficiency issues caused by sellers manually collecting materials are resolved, achieving high-quality and efficient video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ALIBABA INT INTERNET IND CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-05-05
AI Technical Summary
When e-commerce platform sellers generate product videos, current technology relies on manually collecting and uploading materials, resulting in poor video quality and richness, cumbersome operation, and low efficiency.
Design a video middleware platform that directly connects to the data server of an e-commerce platform. The platform sends video generation requests through the seller's client, obtains product metadata and material sets, and automatically selects target materials to generate videos based on the product metadata. It also introduces a reasoning mechanism based on thought chains for intelligent filtering.
It improves the quality and richness of video generation, simplifies the operation process, and increases generation efficiency. Sellers only need to select the target product to generate a high-quality video.
Smart Images

Figure CN121985191A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of video generation technology, and more particularly to a video generation method and system, a video middleware platform, media, and program products. Background Technology
[0002] In today's e-commerce ecosystem, video content has become a key medium for attracting consumers, showcasing product details, and improving conversion rates. For e-commerce sellers, efficiently and cost-effectively producing high-quality product videos is an urgent need to enhance store competitiveness and adapt to the trend of content marketing. Currently, the mainstream methods for sellers to generate product videos typically rely on manual operation and third-party tools. Sellers need to collect or shoot original materials such as product images, video clips, text, and music, and then upload these materials to independent third-party video editing software or online production platforms for video production. However, on the one hand, the amount of materials uploaded by sellers is often relatively small, resulting in poor quality and richness of the generated videos; on the other hand, sellers need to perform a series of operations such as material collection and uploading during the video generation process, making the operation cumbersome and the video generation efficiency low. Summary of the Invention
[0003] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a video generation method is proposed, applied to a video middleware connected to a data server of an e-commerce platform and a seller client of the e-commerce platform, wherein the data server is used to store product metadata and product materials related to products sold on the e-commerce platform; the method includes: Receive the video generation request sent by the seller's client; In response to the video generation request, the system obtains the product metadata and product material set of the target product from the data server, and determines the target product material for generating the video from the product material set based on the product metadata of the target product. The target product material is input into the video generation model, so that the video generation model generates a target video based on the target product material.
[0004] According to a second aspect of one or more embodiments of this specification, a video generation method is proposed, applied to a seller client of an e-commerce platform. The seller client is communicatively connected to a video middleware platform, and the video middleware platform is also communicatively connected to a data server of the e-commerce platform. The data server is used to store product metadata and product materials related to products sold on the e-commerce platform. The method includes: Display the video generation control on the seller backend management page; In response to detecting an operation command for the video generation control, a video generation request is sent to the video platform; The target video returned by the video platform in response to the video generation request is obtained; the target video is generated by the video generation model called by the video platform based on the target product material, the target product material is determined from the product material set of the target product based on the product metadata of the target product, and the product metadata and product material set of the target product are obtained from the data server of the e-commerce platform.
[0005] According to a third aspect of one or more embodiments of this specification, a video middleware platform is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor executes the executable instructions to implement the steps of the method as described in the first aspect of one or more embodiments of this specification.
[0006] According to a fourth aspect of one or more embodiments of this specification, a video generation system is proposed, including an e-commerce platform, a seller client, and a video generation middleware platform; The e-commerce platform includes a data server for storing product metadata and product materials related to the products sold on the e-commerce platform; The seller client is used to send a video generation request to the video platform; The video middleware platform is used to perform the methods described in the first aspect of one or more embodiments of this specification.
[0007] According to a fifth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first or second aspect of one or more embodiments of this specification.
[0008] According to a sixth aspect of one or more embodiments of this specification, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as described in the first or second aspect of one or more embodiments of this specification.
[0009] As described in the above embodiments, this specification designs a video middleware platform to connect with the data server and seller client of an e-commerce platform. Sellers send video generation requests to the video middleware platform through their seller clients. The video middleware platform responds to these requests by retrieving product metadata and a set of product materials from the product database, and then selects target product materials from the set of product materials based on the product metadata to generate the video. On the one hand, since the e-commerce platform's product database stores a large number of rich product materials, generating videos based on these materials effectively improves the quality and richness of the generated videos. On the other hand, in the video generation process described above, sellers only need to perform simple operations such as initiating the video generation process and selecting target products on their seller clients; they do not need to collect and upload materials themselves, making the operation simple and efficient. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of a system architecture provided in an exemplary embodiment.
[0011] Figure 2 This is a flowchart of a video generation method applied to a video middleware platform, provided as an exemplary embodiment.
[0012] Figure 3 This is a schematic diagram of a seller backend management page provided in an exemplary embodiment.
[0013] Figure 4 This is a schematic diagram of an input mode selection page provided in an exemplary embodiment.
[0014] Figure 5 This is a schematic diagram of a product list provided in an exemplary embodiment.
[0015] Figure 6 This is a schematic diagram of a product material selection process provided in an exemplary embodiment.
[0016] Figure 7 This is a schematic diagram illustrating a video attribute information configuration process provided in an exemplary embodiment.
[0017] Figure 8 This is a schematic diagram of a video summary provided in an exemplary embodiment.
[0018] Figure 9 This is a schematic diagram of a video preview process provided in an exemplary embodiment.
[0019] Figure 10 This is a schematic diagram illustrating the overall functionality of a video middleware platform as provided in an exemplary embodiment.
[0020] Figure 11This is a schematic diagram of a system architecture for a video middleware platform provided in an exemplary embodiment.
[0021] Figure 12 This is a schematic diagram of a video upload link provided in an exemplary embodiment.
[0022] Figure 13 This is a flowchart of an exemplary embodiment of a video generation method applied to a seller client.
[0023] Figure 14 This is a schematic diagram of the structure of a device provided in an exemplary embodiment.
[0024] Figure 15 This is a block diagram of a video generation apparatus applied to a video middleware platform, provided as an exemplary embodiment.
[0025] Figure 16 This is a block diagram of a video generation apparatus applied to a seller client, provided as an exemplary embodiment. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0028] In related technologies, sellers on e-commerce platforms need to collect materials themselves when generating product videos, and then upload the collected materials to third-party video editing software or online production platforms for video production. The videos generated in this way are of poor quality and richness, and the operation is cumbersome and inefficient.
[0029] To address the aforementioned issues, this specification proposes a video middleware platform. This platform directly connects to the data server and seller client of an e-commerce platform. Sellers can send video generation requests to the middleware platform via their seller client. The middleware platform responds to these requests by retrieving the seller's published product list, allowing the seller to select a target product. The platform then retrieves product metadata and a set of product materials from the product database, and selects the target product material from the set based on the product metadata to generate the video. On one hand, since the e-commerce platform's product database stores a large quantity and rich content of product materials, generating videos based on these materials effectively improves the quality and richness of the generated videos. On the other hand, in the video generation process, sellers only need to perform simple operations such as initiating the video generation process and selecting the target product on their seller client; they do not need to collect and upload materials themselves, making the operation simple and efficient. The specific implementation of this specification's embodiments is illustrated below with reference to the accompanying drawings.
[0030] Figure 1 This is a schematic diagram of the architecture of a video generation system provided in an exemplary embodiment. For example... Figure 1As shown, the system may include an e-commerce platform 111, a video middleware platform 112, a network 113, and a seller client 114. The e-commerce platform 111 may include several servers, such as data servers, application servers, cache servers, and search servers. The data servers store product information, seller information, and other data related to the e-commerce platform; the application servers handle service requests, such as login requests, order creation requests, and order query requests; the cache servers cache some data from the data servers to alleviate their load and improve response speed; and the search servers provide product search services. Each of these servers can be a physical server containing an independent host, or a virtual server hosted in a host cluster. The video middleware platform 112 may be a core component integrated within the e-commerce platform 111, focusing on the intelligent production and processing of video content, or it may be a dedicated service platform deployed independently of the e-commerce platform 111 and connected and coordinated with it via the network 113. When the video middleware 112 is integrated into the e-commerce platform 111, the video middleware 112 can interact with the seller client 114 via an application server, or it can interact with the seller client 114 directly. The seller client 114 can be a program running on an electronic device, which can be a native application installed on the electronic device, or it can be a mini-program, quick app, or other similar form. The electronic device can include, but is not limited to, personal computers (PCs), mobile phones, tablets, laptops, personal digital assistants (PDAs), wearable devices (such as smart glasses, smartwatches, etc.), etc., and one or more embodiments in this specification do not limit this.
[0031] As for the network 113 that facilitates interaction between the seller client 114, the e-commerce platform 111, and the video conferencing platform 112, communication can be achieved using either wired or wireless networks, based on the communication methods supported by the corresponding electronic devices. This manual does not impose any restrictions on this. For example, a PC can support both wired and wireless communication, so it can use either wired or wireless networks as needed. Mobile phones typically only support wireless communication, so they can use wireless networks for communication.
[0032] Figure 2 This is a flowchart of a video generation method according to an embodiment of this specification. The video generation method can be applied to a video middleware 112 connected to a data server of an e-commerce platform 111 and a seller client 114. The data server is used to store product metadata and product materials related to products sold on the e-commerce platform 111. The method includes: Step S12: Receive the video generation request sent by the seller client 114; Step S14: In response to the video generation request, obtain the product metadata and product material set of the target product from the data server, and determine the target product material for generating the video from the product material set based on the product metadata of the target product; Step S16: Input the target product material into the video generation model so that the video generation model can generate the target video based on the target product material.
[0033] In step S12, the video middleware platform 112 can receive video generation requests sent by the seller client 114. For example... Figure 3 As shown, sellers can access the seller backend management page 100 on the seller client 114. This seller backend management page 100 can embed a video generation control 101. Figure 3 In the example shown, the seller backend management page 100 can be a video management page, used to manage the videos that sellers have published. Sellers can view the list of published videos, the publication time of each video, the number of associated main images, the number of associated product details, publishing suggestions, review status, quality inspection status, video quality, and other information on this video management page. Sellers can also perform operations such as viewing video optimization suggestions, associating product main images, and associating product details on published videos. In addition to the video management page, the seller backend management page 100 may also include other functional pages, including but not limited to, those for managing products, orders, customers, marketing activities, or store settings.
[0034] When a seller interacts with the video generation control 101 (e.g., clicks), the seller client 114 sends a video generation request to the video platform 112. This request may include seller identification information (seller ID) to identify the seller. The seller can pre-login to the seller client 114. After successful login, the e-commerce platform 111's application server can return the seller ID to the seller client 114, which can then store it. After detecting an interaction with the food generation control, the seller client 114 can include the seller ID as a request parameter when constructing the food generation request, sending it along with the video generation request to the video platform 112.
[0035] In step S14, after the seller successfully publishes a product, the e-commerce platform 111 will persistently store the product information (including product metadata and product materials) in the product database and establish a unique identifier (product ID) for the product associated with the seller's seller ID. After receiving a video generation request from the seller client 114, the video middleware 112 can send a query request to the data server to retrieve the product ID associated with the seller ID. Based on the retrieved product ID, it requests the corresponding product information from the data server and returns the product information returned by the data server to the seller client 114, so that the seller client 114 can render a product list on the user interface based on the product information returned by the video middleware 112.
[0036] Sellers can perform a selection operation on the seller client to choose a target product from the product list. The seller client 114 can send the target product's product ID to the video platform 112, enabling the video platform 112 to retrieve the target product's metadata and a set of product materials from the data server based on the product ID. The product metadata can be multimodal data, including but not limited to at least one of the following: product title, product image, product category, product attributes (such as color, material, weight, size, etc.), product price, and product interaction data (such as click count, favorite count, add-to-cart count, purchase count, etc.). The set of product materials can include multiple product materials, including but not limited to at least one of the following: main product image, product detail image, product scene image, product qualification image, product close-up image, and 3D model of the product.
[0037] Furthermore, the video platform 112 can provide multiple information input modes, such as automatic information input mode, URL input mode, and manual input mode. The automatic information input mode refers to the video platform 112 obtaining a product list from the data server, the seller client 114 selecting the target product from the product list, and then the video platform 112 automatically retrieving the product information from the data server, which is the method described in the previous embodiments. The URL input mode refers to the seller client 114 entering a URL and retrieving product information from the corresponding page based on the entered URL. The manual input mode refers to the seller manually entering product information on the seller client 114. After the seller performs an operation on the video generation control 101 on the seller client 114, the seller client 114 can display an input mode selection page 200, which includes selection components corresponding to each information input mode. Figure 4As shown, the selection components include selection component 201 for automatic information input mode, selection component 202 for URL input mode, and selection component 203 for manual input mode. After the seller operates selection component 201 for automatic information input mode, product information can be obtained through the process described in the aforementioned embodiment.
[0038] After the seller operates the selection component 201 corresponding to the automatic information input mode, the list of products published by the seller can be displayed on the seller client 114. For example... Figure 5 As shown, the product list includes items such as dolls, towels, and pajamas. In some embodiments, the products in the product list can be divided into one or more groups. Sellers can view the products in each group through the grouping control on the seller client 114 and select the target product from any group.
[0039] After the seller selects a target product, the seller client 114 can send a selection command to the video platform 112. This selection command can include the product ID of the target product. The video platform can then determine the target product based on the product ID in the selection command and retrieve the product metadata and product material set of the target product from the data server.
[0040] Since not all acquired product materials are suitable for video generation, the video platform can filter the product materials after acquiring the set to select the target product materials suitable for video generation. In some embodiments, different target products are suitable for different target product materials. For example, clothing products need to showcase styling effects, so product images including models wearing the clothing and multi-angle views can be prioritized as target product materials; 3C digital products need to highlight functional features and a sense of technology, so product materials including function demonstrations, interface close-ups, and technical parameter comparisons can be prioritized as target product materials; home furnishing products need to reflect design and spatial coordination, so materials including scene-based displays, size comparisons, and material details can be prioritized as target product materials. Therefore, target product materials can be determined from the product material set based on the target product's metadata, thereby automatically selecting the most relevant target product materials that best showcase the core selling points of different types of target products, thus improving the targeting and display effect of the generated target video.
[0041] In some embodiments, the operation of determining the target product material can be performed by a target inference model. The target inference model can be a large language model. To overcome the limitations of traditional keyword matching or single instructions in complex material selection scenarios, and to improve the accuracy, interpretability, and adaptability of the selection process, this embodiment introduces a thought chain-based inference mechanism. The thought chain-based inference mechanism is illustrated below with examples.
[0042] First, pre-generated thought chain prompts can be obtained. These prompts define multi-level reasoning steps for filtering product materials based on product metadata. The output of each level serves as the input to at least one subsequent reasoning step, thus forming a structured, multi-step reasoning logic framework. This framework decomposes the complex task of material selection into a series of interconnected and progressively advancing reasoning steps. In some embodiments, the multi-level reasoning steps in the thought chain prompts may include: (i) Obtain the task background information for the video generation task. This background information may include, but is not limited to, the selection criteria for product materials, the parsing process for product materials, the first evaluation criterion for assessing the match between product materials and product metadata, user information of the video audience, and the second evaluation criterion for assessing the match between product materials and the video audience. The selection criteria for product materials define the characteristics of eligible product materials, such as resolution and the proportion of the product subject in image-based product materials. The parsing process for product materials defines a series of processes in the material selection process, such as material quality filtering, material parsing, material-product matching, material-audience matching, material weight determination, and target material selection. The first evaluation criterion establishes quantitative or logical comparison rules between the material content and the inherent attributes of the product. For example, it assesses whether the category of the product displayed in the product material matches the "category" in the product metadata, whether the color of the product in the product material matches the "color" attribute in the product metadata, whether the style of the product material (e.g., minimalist, luxurious) matches the "style" description in the product metadata, and whether the product material clearly displays the "core selling points" (e.g., waterproof, foldable) in the product metadata. This standard directly determines the accuracy and fidelity of the product materials' representation of the product itself. The second evaluation standard measures the degree to which the content of the product materials resonates with the video audience on an emotional, cultural, or consumer psychology level. For example, it assesses whether the life scenarios presented in the product materials (such as commuting to work, outdoor sports, and family time) match the "typical scenarios" in the video audience's user profile; whether the emotional tone conveyed by the product materials (such as warm, cool, and professional) aligns with the video audience's "preference style"; and whether the characters appearing in the product materials (such as age and clothing) easily evoke "identity recognition" and resonance from the video audience. This standard aims to enhance the emotional appeal and conversion potential of the generated videos.
[0043] (ii) Material quality filtering, which involves selecting multiple candidate product materials that meet the filtering criteria from the product material set. This step serves as a prerequisite for subsequent intelligent inference and aims to quickly eliminate invalid or low-quality product materials through hard rules, ensuring that all product materials input into the target inference model meet the minimum usability standards, including but not limited to: selecting product materials with a resolution not lower than a preset resolution threshold and a proportion of the product subject in the product materials not lower than a preset proportion threshold.
[0044] (iii) Material Analysis: This involves analyzing each candidate product material selected according to the analysis process to determine its material information. Specifically, this analysis process may include: 1) Subject Object Analysis: Identifying and labeling the core subjects (such as specific products, human models, usage scenarios) in the product materials and their positions, postures, or interaction relationships; 2) Visual Feature Analysis: Extracting attributes such as color composition, texture, visual style (such as minimalist, luxurious), and image composition of the product materials; 3) Text Information Analysis: Detecting and identifying embedded text or watermarks that may be contained in the materials and converting them into processable text data; 4) Scene and Emotion Analysis: Determining the type of scene depicted by the materials (such as indoor, outdoor, office) and the emotional tone it conveys (such as warm, professional, dynamic). Through the above analysis, the original unstructured multimedia materials are transformed into semantically rich structured descriptive information, providing a data foundation for subsequent accurate matching and evaluation based on product metadata and user profiles that can be directly calculated and reasoned.
[0045] (iv) Material and Product Matching: This involves evaluating the first degree of matching between the material information of candidate product materials and the product metadata of the target product based on the first evaluation criteria. The target inference model can compare the product metadata and the material information of candidate product materials line by line according to the first evaluation criteria (e.g., "evaluating whether the product materials accurately display the product entity, color, style, and core functions"). For example, "The main body of the material is a thermos cup, which matches the product category; the color description is 'pink,' which highly semantically matches the metadata 'cherry blossom pink'; the style description is 'simple,' which is consistent with 'simple and stylish'; however, the material does not directly display the APP temperature control function, so the match is incomplete on this selling point." Then, a quantitative first degree of matching score is output based on the comparison results.
[0046] (v) Material and Audience Matching: This step assesses the second degree of matching between the material information of candidate products and the user information of the video audience based on the second evaluation criteria. In this step, the target inference model compares the material information with the user information of the video audience (e.g., "User group: young urban white-collar women; preferences: refined, healthy, and technological lifestyle") and infers based on the second evaluation criteria (e.g., "Assess whether the aesthetic style, scene, and emotional communication of the material match the preferences of the target user group"). For example, "The material scene is 'office desk,' which precisely matches the 'white-collar' professional scene; the overall 'simple' style matches the group's preference for 'refined'; however, the material is weak in conveying the emotional message of 'technological' and 'healthy lifestyle'." Based on this, the target inference model outputs another independent second degree of matching score.
[0047] (vi) Material weighting: For each candidate product material, its weight is determined based on its first and second matching degrees. For example, a preset importance coefficient can be assigned to the first and second matching degrees, indicating whether the material selection process focuses more on the faithful representation of the product itself or its appeal to the video audience. Subsequently, the first and second matching degrees are weighted based on the aforementioned importance coefficients to obtain the weight of the candidate product material.
[0048] (vii) Target material selection, which is to select the target product material from multiple candidate product materials based on their weights. For example, candidate product materials with a weight greater than a preset value can be selected as the target product material, or candidate product materials with the top K weights (K is a positive integer) can be selected as the target product material.
[0049] Then, the thought chain prompts, the target product's metadata, and the set of product materials can be input into the target inference model. This allows the model to perform thought chain reasoning on the target product's metadata and material set according to multi-level reasoning steps, and determine the target product materials based on the reasoning results. Through this method, the target inference model no longer performs end-to-end black-box decision-making, but is guided to follow predetermined thought chain steps, performing step-by-step reasoning with visible intermediate results. It first understands the semantics of the product metadata, then progressively compares, evaluates, and logically judges the selection criteria and product material characteristics in multiple rounds, ultimately outputting the target product materials after rational reasoning. This process significantly improves the accuracy of product material selection, product context matching, and user preference matching, and makes the model's decision-making process more transparent and controllable.
[0050] In more specific examples, thought chain templates with specified data formats can be generated based on multi-level reasoning steps. Thought chain templates can adopt a structure of fixed frame + variable placeholders: Fixed content is used to describe the process framework, output format requirements, and evaluation logic of the reasoning steps; Variable placeholders can be marked with specific symbols (such as "{}") and used to populate specific information at execution time.
[0051] In addition, the output data format of the target inference model (such as JSON format) can be defined in the thought chain template.
[0052] Some implementation examples of mind chain templates can be in the following Markdown format: # Product Material Selection Mindset ## System Role Definition {System Role Description} ## Task Background A video of type {video} needs to be created for {product name}, with the target audience being {user group} and the video's purpose being {video purpose}.
[0053] ## Available Material Library Information The resource library contains {total number of resources} candidate resources, and the types include {a list of resource types}.
[0054] ## Multi-level reasoning steps ### Step 1: Understanding the Task Background and Analyzing the Standards Input information: - Basic Product Information: {Basic Product Information} - Key selling points of the product: {List of product selling points} - Video production goals: {Video production goals} - Material selection criteria: {Description of selection criteria} - Material parsing process: {Description of parsing process} - First Evaluation Criterion (Product Matching): {Details of First Evaluation Criterion} - Second evaluation criterion (user matching degree): {Details of the second evaluation criterion} - User profile information: {User profile details} Task instructions: Based on the above information, please understand the core requirements of this video production task and identify the key dimensions for selection.
[0055] Output format requirements: json { "Task Summary": "A concise task description", "Filter Priority Sort": ["Highest Priority Dimension", "Second Highest Priority Dimension", ...], "Key Matching Point": { "Product Side": ["Required Product Feature 1", "Product Feature 2", ...], "User side": ["User point 1 that must be attracted", "User point 2", ...] }, "Filtering Constraints": ["Condition 1 that must be met", "Condition 2", ...], "Evaluation weight recommendations": { "Product Matching Weight": Suggested value, "User Matching Weight": Suggested Value } } After generating the thought chain template, the interface between the thought chain template and the target inference model can be called to input the thought chain template, product metadata, and product material set into the target inference model. The target inference model can parse the template structure, identify the fixed framework and variable placeholders, and then automatically fill the corresponding placeholders with product metadata (such as product name, selling points, specifications, etc.) and product material set (such as descriptions of images and videos), forming a specific and executable prompt sequence. Next, the target inference model can execute the inference sequentially according to the multi-level inference steps preset in the thought chain template. Each step of inference is based on the output results of the previous step and the original input data, progressing step by step: from understanding the task background and screening candidate materials, to deeply analyzing material information, to evaluating product matching degree and user matching degree, and finally calculating the comprehensive weight and determining the final target material. In this process, the target inference model can cache the intermediate results of each step to ensure accurate information transfer between different steps, while maintaining the coherence of the inference logic. After completing all steps, the target inference model can output a structured final result, including recommended target materials, detailed recommendation reasons, alternative solutions, and a summary of the entire inference process. Additionally, an interpretability report can be generated, recording the complete reasoning process for subsequent review and optimization.
[0056] The above embodiments utilize the massive collection of product materials accumulated on the e-commerce platform 111 for video generation, breaking through the bottleneck of traditional solutions that rely on limited user-uploaded materials. This significantly improves the richness, visual diversity, and professional quality of the generated videos, effectively avoiding the problem of generating only monotonous, awkwardly connected slideshow-like videos. At the same time, precisely because the platform materials used are vast and diverse, they inevitably contain materials of varying quality or those mismatched with the current video generation goals. Therefore, an intelligent and automated material filtering mechanism is introduced, enabling the video platform 112 to accurately and efficiently identify and combine target product materials that are highly relevant to the target product, of high quality, and aligned with the preferences of the video's target audience from a massive pool of candidates. Ultimately, this achieves a balance between the richness, professionalism, and user targeting of the video content.
[0057] The training process of the target reasoning model is illustrated below with an example.
[0058] First, sample data can be obtained, including sample thought chain prompts, sample product metadata, sample product materials, and tag information. The sample thought chain prompts include multi-level inference steps and the labeled output results corresponding to each level of inference. The tag information is used to indicate whether the sample product materials are selected as materials for generating videos. The multi-level inference steps in the sample thought chain prompts can be referred to the aforementioned embodiment, and will not be repeated here. The labeled output results corresponding to each level of inference can be obtained through manual annotation. For example, in the "material analysis" step, the accurate subject, color, and style description of the sample product materials can be labeled, and this annotation information serves as a reference for the subsequent output content of the base model. The tag information serves as the final supervision signal, clearly indicating whether the sample product materials should be selected as positive sample materials for generating videos in real-world scenarios.
[0059] Then, sample data can be input into a base model (such as a general large language model) to enable the base model to perform reasoning on the sample product metadata and sample product materials based on the multi-level reasoning steps in the sample thought chain prompts, obtaining the reasoning results corresponding to each level of reasoning. During training, the base model is required to strictly follow the multi-level steps defined in the sample thought chain prompts, performing step-by-step reasoning on the input sample product metadata and sample product materials. The base model needs to generate intermediate outputs for each reasoning step sequentially. For example, first, it generates a quality judgment on the sample product materials, then parses the material information, then provides the evaluation reasons and scores for the first and second matching degrees, and finally gives a conclusion on whether to select the material. Here, sample product metadata and sample product materials are the product metadata and product materials of the sample product, respectively. Sample products can be determined based on the interaction data of products on e-commerce platforms. For example, the A / B test conversion rates of videos corresponding to multiple products on e-commerce platforms can be obtained, and sample products can be determined from multiple products based on these A / B test conversion rates. For example, select products whose videos generated in recent A / B tests significantly improved purchase conversion rates or key interaction metrics (such as click-through rate and completion rate) as sample products; or select products whose videos have been verified through A / B testing and whose conversion rates are significantly higher than the average level of the same category as sample products.
[0060] Next, the loss for each inference step is determined based on the difference between the inference results at each level of the base model and the labeled output results at each level. The total loss is then determined based on the losses at each level of the inference steps, and the base model is trained using this total loss to obtain the target inference model. In this step, the local loss for each step is obtained by calculating the difference between the inference results generated by the base model at each level and the labeled output results for that step (e.g., using loss functions such as cross-entropy). Subsequently, the local losses of all steps are weighted and fused with the final decision loss (i.e., the difference between the model's final selection conclusion and the label information) to calculate the total loss. This method ensures that the model learns interpretable and generalizable complex selection capabilities by forcing the base model to learn the correct inference logic at each step, rather than just the final output answer. Finally, the parameters of the base model are iteratively updated based on the calculated total loss using the backpropagation algorithm. After multiple rounds of training on a large number of samples, a target inference model capable of accurately understanding and executing the product material selection thought process can be obtained. This model not only mastered the final screening decision, but also internalized a complete and reasonable reasoning path from quality judgment to multidimensional matching evaluation.
[0061] In some embodiments, to enable the training process of the target inference model to dynamically adapt to the service objectives of the e-commerce platform (such as improving click-through rate and purchase conversion rate), the loss calculation can be closely integrated with real-time interaction data of the product. Specifically, the product metadata can include dynamic interaction data extracted from user behavior, such as the product's recent click-through rate, add-to-cart rate, purchase conversion rate, average viewing time of the corresponding video, and interaction rate (such as likes, comments, and shares). This interaction data can be obtained in real time from the e-commerce platform's real-time interaction logs and stored on the platform's data server, forming product behavior characteristics that evolve over time.
[0062] Based on this type of interactive data, for each sample product, a weighting factor can be introduced to weight the total loss corresponding to that training sample, resulting in a weighted total loss for that sample product. In some embodiments, a first weighting factor can be determined based on the click-through rate of the sample product, a second weighting factor based on the purchase conversion rate of the sample product, and a third weighting factor based on the product viewing time of the sample product. The first, second, and third weighting factors are then weighted to obtain a total weighting factor. The total loss is then weighted based on this total weighting factor to obtain a weighted total loss. The parameters of the base model are then iteratively updated based on the weighted total loss. By adjusting the proportions of different weighting factors, the inference model can be flexibly guided to prioritize learning the most critical service objectives at the current stage (e.g., initially focusing on attracting clicks, and later focusing on conversion achievement), thereby improving training efficiency and the final service adaptability of the inference model. By weighting the total loss and then using the weighted total loss to adjust the parameters of the inference model, the parameter adjustment results can be more influenced by sample products that have shown clear positive user feedback (such as high click-through rate and high conversion rate) in historical A / B tests. This makes the gradient descent direction of the inference model more strongly point to the feature space associated with these high feedback samples when updating parameters.
[0063] like Figure 6 As shown, after filtering out the target product materials, the product metadata (such as product title and product description information) and the filtering results of the target product materials can be sent to the seller client 114 for display. The seller client 114 can display the collection of product materials on its user interface and identify the target product materials filtered by the video platform 112 through methods such as highlighting borders, selection marks, or pinning to the top. Through this embodiment, the seller can be intuitively presented with a complete picture of the available materials and the specific results of the system's intelligent filtering.
[0064] In some embodiments, the target product materials can also be adjusted in response to adjustment instructions from the seller client 114. For example, a user can delete the selected target product material, select the selected product material as the target product material, or manually upload other target product materials on the seller client 114.
[0065] In some embodiments, the video attribute information of the target video can also be configured through the seller client 114. Specifically, the video platform 112 can obtain the video attribute information sent by the seller client 114, and input the video attribute information and the target product material into the video generation model, so that the video generation model generates a target video with the video attribute information based on the target product material.
[0066] Specifically, the video attribute information can be first encoded into a conditional vector using a text encoder, and the target product material can be encoded into a visual feature tensor using a visual encoder. Then, the conditional vector, visual feature tensor, and a randomly initialized noisy video tensor are input into the denoising network of the diffusion model. This denoising network adopts a U-Net architecture, mainly consisting of an encoder and a decoder. The encoder contains multiple cascaded residual blocks, each mainly composed of a cross-attention module and a spatiotemporal transformation module. The cross-attention module takes the conditional vector and visual feature tensor as input, aligns and fuses the conditional information represented by the conditional vector into the visual context represented by the visual feature tensor through an attention mechanism, outputting a conditionally enhanced visual guidance feature. The spatiotemporal transformation module takes the aforementioned visual guidance feature and the hidden layer representation of the current noisy video as input, first extracting local spatiotemporal features through 3D convolution, and then modeling long-range dependencies in the spatial dimension (between pixels within the same frame) and the temporal dimension (between frames at the same position) through a separate spatiotemporal self-attention mechanism, finally outputting abstracted and compressed deep features. The decoder is symmetrical to the encoder and consists of multiple upsampled residual blocks. Each block takes detailed features from the corresponding layer of the encoder (the corresponding network layers in the encoder and decoder are interconnected via skip connections) and high-level semantic features from deeper layers as input. First, it is upsampled through 3D transposed convolution to restore spatial resolution. Then, using lightweight spatiotemporal convolution and attention mechanisms, the detailed and high-level semantic features transmitted by the encoder are fused and calibrated, gradually refining the features and reconstructing a detailed and spatiotemporally coherent sequence of video frames. Finally, the decoder output is mapped through a projection layer to obtain the predicted denoised video data.
[0067] like Figure 7As shown, a configuration page can be displayed on the seller client 114, where sellers can configure video attribute information. Video attribute information may include, but is not limited to, at least one of the following: video style (e.g., ink painting style, 3D style, animation style, etc.), video aspect ratio, video language, digital humans in the video, and the video's target audience. After generating the target video, the seller may distribute it to external platforms. Different platforms may have different video styles; therefore, sellers can select different video styles by choosing the platform for the target video. Furthermore, users can randomly select digital humans, customize digital humans, or choose not to add digital humans to the target video. When randomly selecting digital humans, the video platform 112 can obtain the attribute information (e.g., gender, hairstyle, skin color, age, etc.) of the digital humans in the most recent N target videos generated for the seller, and determine the attribute information of the digital humans in the target product generated this time based on the attribute information of the digital humans in the most recent N target videos.
[0068] In step S18, a target video can be generated by the video generation model based on the target product materials. The generated target video can be stored in the object storage service of the video platform 12, or it can be directly sent to the seller client 114. In some embodiments, a two-stage video generation process of video summary generation + video content generation can be adopted. An example is given below.
[0069] In the first stage, a video summary can be generated based on the target product's metadata. This video summary generation process can be performed by a summary generation model (e.g., a large language model). The summary generation model takes the target product's metadata (such as title, category, attributes, selling points list, video audience information, etc.) as input, and through understanding, summarizing, and creatively reorganizing it, outputs a structured video summary. This video summary is used to describe the video content and style, and may include, but is not limited to, the following information: Narrative structure and shot sequence: For example, opening close-up of the product's core selling points → panoramic display of usage scenarios → contrast to highlight material details → ending with the presentation of the brand logo.
[0070] Visual and copywriting guidelines: For example, emphasizing product features and defining the main advertising slogan.
[0071] Pace and duration suggestions: For example, the overall pace should be brisk, and the duration should be kept within 15 seconds.
[0072] Audio and style requirements: For example, the background music should be modern electronic music, and the overall style should lean towards technology and fashion.
[0073] The video conferencing platform 112 can send the generated video summary to the seller client 114 for display. Some embodiments of the video summary include... Figure 8 As shown. In this embodiment, the summary generation model can generate several candidate video summaries, and different candidate video summaries can correspond to different styles (such as positive encouragement, lighthearted and cheerful, etc.). Users can select one of the multiple candidate video summaries for subsequent video generation, or they can instruct the summary generation model to regenerate the video summary through the "Regenerate" control on the interface. In addition, sellers can also send modification instructions for the video summary to the video platform 112 through the seller client 114. The video platform 112 can respond to the modification instructions received from the seller client 114 and modify the video summary accordingly.
[0074] In the second stage, video summaries and target product materials can be input into the video generation model, enabling the model to generate the target video based on these materials. The model first parses the video summary to understand its narrative instructions and stylistic requirements. Then, based on the sequence of shots in the summary, it automatically sorts, crops, and colors the target product materials, generating appropriate transition effects. It can also automatically generate and synthesize subtitles, title animations, or voice-overs based on the text guidance in the video summary; it can also automatically match and insert suitable background music, sound effects, and motion graphics elements according to stylistic requirements; furthermore, it can add digital humans to the video. After completing the temporal alignment and spatial compositing of all elements, the video generation model renders and outputs the final target video.
[0075] In some embodiments, see Figure 9 The video generation model can first generate multiple video preview files based on the target product materials, and then send these multiple video preview files to the seller client 114 for display. In response to receiving the target video preview file (such as...) from the seller client 114 among the multiple video preview files... Figure 9 The selection command for video preview file 1) generates a target video with the visual characteristics corresponding to the target video preview file. The video preview file can be a low frame rate preview sequence composed of multiple sequentially arranged images with specific visual characteristics. This sequence can simulate and present the core visual style and content overview of the target video. The images in the sequence can be generated using lower resolution and compression rates, significantly reducing the computational overhead and generation time required for real-time rendering, facilitating the rapid provision of multiple visual options for sellers to choose from and confirm. Different video preview files correspond to different visual configuration parameters, such as different positions and image attributes of the digital human in the frame, different aspect ratios of the video frame, and / or different color tones and visual filters applied.
[0076] In some embodiments, the video platform 112 can support multiple video generation modes. Different video generation modes are used to generate different types of target videos, and different types of target videos are generated by different video generation models. The type of target video can be determined based on its function, such as product promotion videos, product function demonstration videos, and product marketing videos. Sellers can select a video generation mode on their seller client 114. The video platform 112 can obtain the target video generation mode selected by the seller client 114 from the multiple video generation modes, and input the target product materials into the video generation model corresponding to the target video generation mode, so that the video generation model corresponding to the target video generation mode generates the target video based on the target product materials.
[0077] In some embodiments, the video middleware 112 may also perform at least one of the following operations: The target video is translated. The video platform 112 can perform multilingual translation and adaptation of the target video's audio track, subtitles, or embedded text elements. For example, it can automatically identify the original speech or text in the video, translate it into the language of the target distribution area, and generate corresponding dubbing audio tracks or subtitle files, thereby achieving rapid localization of video content to meet cross-regional distribution needs.
[0078] Video splitting of target videos. Video platform 112 can automatically generate multiple derivative video versions with differences in content, length or format from a main video template through intelligent editing, image recombination, replacement of local elements (such as background, subtitles, product prices) or overlay of different channel logos, in order to meet the needs of differentiated delivery on different advertising platforms.
[0079] Publish the target video to the designated platform. The video platform 112 can be pre-configured to connect to the product details page of e-commerce platform 111 or the publishing interface of external advertising platforms. After the video is rendered, the target video and related information can be automatically published to one or more designated locations according to rules or after seller confirmation.
[0080] The target video is transcoded. The video platform 112 can automatically convert the video into multiple versions with different encoding formats (such as H.264 / AVC, H.265 / HEVC), resolutions (such as 4K, 1080p, 720p), and bitrates according to the format specifications of each publishing platform, network conditions, and terminal device performance requirements, to ensure that the video achieves the best playback experience and compatibility on the widest range of devices.
[0081] Risk detection is performed on target videos. This includes, but is not limited to, using image recognition, audio analysis, and text detection technologies to scan video footage, audio, and subtitles to identify whether they contain infringing content, prohibited items, sensitive information, or risk elements that do not comply with the platform's publishing guidelines. This process provides review prompts to operations personnel or automatically executes blocking actions to ensure content security and compliance.
[0082] In some embodiments, the video platform 112 also provides a hosting function. This function enables the system to automatically complete the entire process from identifying needs to generating, processing, and publishing target videos based on preset rules or external events, without requiring manual triggering by the seller. Specifically, when a new product is successfully published on the data server of the e-commerce platform 111, this listing event can be monitored or detected by the hosting service of the video platform 112. Subsequently, the hosting service will automatically use the currently published new product as the target product and execute the video generation process described in the aforementioned embodiments. Through this hosting function, the video platform 112 achieves an upgrade from "passively responding to seller operations" to "actively responding to triggering events." It ensures the timeliness, consistency, and professionalism of video content generation at important event nodes (such as new product listings), significantly improving the automation level of e-commerce operations and video generation efficiency.
[0083] Figure 10 The overall functionality of the Video Platform 112 is demonstrated, including three modules: content creation, automatic sharing of content across multiple channels, and content delivery performance tracking. In terms of content creation, the Video Platform 112 not only supports video generation but also provides image generation based on templates or text commands, as well as automatic copywriting capabilities. Regarding content generation and traffic acquisition, the Video Platform 112 supports automatic or one-click publishing of generated target videos to the e-commerce platform 111 (such as associated product detail pages) and designated external social media platforms, and provides unified management capabilities for distribution channel accounts, simplifying cross-platform operations. In terms of content delivery performance tracking, the Video Platform 112 can integrate delivery data from various channels, providing visualized analysis of core metrics such as clicks, impressions, and conversions. It can also provide optimization suggestions regarding publication time and / or content type, and offer benchmark references such as industry average data to help sellers objectively evaluate content performance. The entire process—from material selection, video generation model selection, target video generation, video uploading to OSS, video storage, video risk control, and video quality review—can be encapsulated into an Agent service for batch invocation by external access parties (seller client 114).
[0084] The system architecture of the video middleware platform 112 is as follows: Figure 11 As shown. This architecture, from bottom to top, mainly includes the following layers and core modules: Infrastructure Layer: As the cornerstone of the system, it provides global support capabilities. Specifically, the infrastructure layer includes a configuration center, responsible for the centralized management and real-time distribution of service provider weights, dynamic parameters, API keys, etc., enabling policy adjustments without requiring service restarts; and a monitoring center, integrating log tracking, metric collection, and monitoring systems to ensure service observability, security, and stable operation. Service Layer: This includes the entry point layer, which serves as the unified receiving and distribution point for video generation requests. Video generation requests can be HTTP POST requests, including but not limited to some or all of the parameters such as Seller ID, Product ID, Video Generation Task ID, Video Style, and Video Duration. Seller client 114 can send the video generation request to the TV platform's application server. The application server can then forward the request to the video middleware platform 112 through the entry point layer, where preliminary verification and formatting are performed.
[0085] Service Abstraction Layer: As the core of the architecture's intelligent scheduling, it decouples service logic from specific service providers. Specifically, the Service Abstraction Layer includes: a strategy selection module, which intelligently selects a target service provider based on the service provider's weight, cost, real-time load, or performance metrics from the configuration center, and then calls the target service provider's video generation model to generate video; a fault tolerance / degradation module, with built-in retry, circuit breaking, and degradation mechanisms. When the preferred service provider fails, it automatically switches to the backup service provider's video generation model according to the strategy or returns a fallback result, ensuring that the core process is not interrupted; and a load balancing module, which reasonably distributes requests to different nodes when making calls to the same service provider.
[0086] Adapter Layer: Responsible for the specific interface with heterogeneous external services. Specifically, it can provide an independent adapter for each integrated third-party video generation service (as shown in the diagram, vendor adapters A / B / C). Each adapter encapsulates the service-specific parameter adaptation, protocol conversion, and authentication logic, presenting a unified interface to the upper layers.
[0087] When a video generation request arrives at the entry layer, the service selection strategy module of the service abstraction layer routes the request to the target vendor adapter (let's say vendor adapter B) based on the current configuration. Subsequently, the video generation request is passed to vendor adapter B via the fault tolerance / degradation module. This adapter adapts parameters and performs protocol conversion according to vendor B's API specifications, and initiates an external call carrying the API key obtained from the configuration center. Key logs and performance metrics throughout the process are reported to the monitoring center. If the call to vendor B fails and reaches a threshold, the fault tolerance / degradation module will automatically trigger degradation, potentially switching to vendor adapter A to continue executing the request. This architecture, through layered abstraction and policy-based scheduling, enables the video platform 112 to integrate and manage multiple video generation services in a unified, stable, and flexible manner. It achieves seamless vendor switching and elastic scaling, reduces the system's dependence on a single vendor, and significantly improves the professionalism, availability, and maintainability of the entire video generation service through unified monitoring, fault tolerance, and configuration management.
[0088] In some embodiments, the video platform 112, to ensure the compliance, playback compatibility, and system stability of all video content, has designed and implemented a unified and robust video upload and processing chain. For example... Figure 12 As shown, this link serves as the core pipeline, through which all video content (whether it's original material uploaded locally by sellers or finished video products generated by AI services) must be processed, ensuring standardization and controllability of the process. This link consists of the following core modules, collectively completing a closed loop from access to distribution: Video Upload: Provides a unified video upload service portal, accepting video files from various sources. After a successful upload, the system persistently stores the original video file in a unified object storage service and generates a globally unique file identifier, laying the foundation for subsequent processing.
[0089] Video transcoding: To adapt to playback requirements under different terminals and network environments, the stored original videos undergo standardized transcoding processing. Using preset transcoding templates, video streams with various resolutions, bitrates, and container formats (such as HLS and MP4) are generated. This allows for flexible selection of the most suitable version based on specific scenarios (such as product detail pages or mobile news feeds), optimizing user experience and saving bandwidth.
[0090] Risk Control Review: After transcoding, the video will automatically enter a unified content security review process. A multimodal review model can be invoked to simultaneously detect the video's visuals, audio, and subtitles, identifying whether it contains pornographic, politically sensitive, or violent content. If the review fails, the system will immediately notify the uploading service provider via an asynchronous callback mechanism and automatically remove the video to ensure that non-compliant content is not distributed.
[0091] Content storage and quality rating include cloud storage services and content quality review services. Cloud storage services provide storage space for video files with service attributes (such as associated stores and product IDs), enabling resource isolation and management. The content quality review service, based on risk control, automatically rates the technical quality and content experience of videos. Video quality can be categorized into unrated, low-quality, average, high-quality, and premium levels based on multi-dimensional indicators. For "low-quality" videos, the system will further add specific tags, such as: black borders, shaky footage, PPT-style slideshow, excessively low frame rate, abnormal file size, unreasonable video length, substandard resolution, inappropriate aspect ratio, containing contact information, no audio or effective subtitles, completely silent, presence of non-platform watermarks, or low-quality template-generated content (such as Aliwood-style videos). This rating can provide data support for video priority recommendations, traffic allocation, and optimization suggestions for sellers.
[0092] Processing Log Recording: To enhance the transparency and observability of the process, after generating and processing the target video, the video platform 112 can not only return the final access address of the target video to the seller client 114, or directly send the target video to the seller client 114, but also generate and return a structured log file. This log file records key node information throughout the entire video generation process, such as the matching score and final selection confidence of each candidate material in the material selection stage, the reasons for selecting the video generation model and its calling status, the risk control review result code, and the video quality rating score and specific tags. This log provides external access parties with quantifiable and analyzable technical evidence, in addition to the final video file, facilitating effect attribution, problem investigation, and subsequent optimization.
[0093] Stability Guarantee: To ensure high availability and service quality of the link, the system has a built-in multi-layered stability guarantee mechanism, including but not limited to: Multi-service provider mechanism: For dependent external services (such as transcoding and auditing), multiple alternative service providers are integrated. The system monitors the API error rate and response latency of each service provider in real time. When the metrics exceed preset thresholds, traffic is automatically switched to a healthy service provider. Calls between service providers are completely isolated to prevent the spread of failures.
[0094] QPS rate limiting: Implement a tiered query per second rate limiting strategy at the service entry point and when calling various service providers, respectively limiting the call frequency to a single service provider and the total number of global requests to prevent traffic surges from overwhelming the system.
[0095] Exception retry and fallback: When a service call encounters a retryable exception, the system automatically retryes a limited number of times according to the policy. If the call still fails after retrying, the preset fallback logic is executed, such as returning a degraded result or a friendly message, to ensure that the core process is not interrupted.
[0096] The above embodiments, by constructing an automated pipeline that integrates uploading, transcoding, risk control, quality rating, and stability assurance, ensure that the video platform 112 can process massive amounts of video content in a safe, reliable, and efficient manner, providing a solid content infrastructure capability for upper-layer e-commerce services.
[0097] See Figure 13 This specification also provides a video generation method applied to a seller client 114 of an e-commerce platform 111. The seller client 114 is communicatively connected to a video middleware platform 112, which is also communicatively connected to a data server of the e-commerce platform 111. The data server is used to store product metadata and product materials related to the products sold on the e-commerce platform 111. The method includes: Step S22: Display the video generation control on the seller backend management page; Step S24: In response to detecting an operation command for the video generation control, a video generation request is sent to the video platform 112; the video generation request carries seller identification information; Step S26: Obtain the product list returned by the video middle platform 112 in response to the video generation request. This product list is the list of products published by the target seller to which the seller identification information belongs on the e-commerce platform 111. Step S28: In response to detecting a selection operation for a target product in the product list, the identification information of the target product is sent to the video middle platform 112, so that the video middle platform 112 obtains the product metadata and product material set of the target product from the data server, determines the target product material for generating the video from the product material set based on the product metadata of the target product, and inputs the target product material into the video generation model, so that the video generation model generates the target video based on the target product material.
[0098] The specific process executed by the seller client 114 is detailed in the aforementioned embodiments and will not be repeated here.
[0099] In some embodiments, the video generation control displayed on the seller backend management page has a specific style, and the style of the video generation control can change dynamically. For example, the style of the video generation control can be dynamically adjusted based on the seller's historical operation data. Specifically, the seller's static attributes and dynamic behavior data can be continuously collected and analyzed. Static attributes include their store level, main category, historical frequency of video generation, preferred video style, final quality rating of generated videos (such as the percentage of high-quality videos), and performance in A / B testing (such as click-through rate and conversion rate). Based on the analysis of this data, the strategy decision engine within the system can identify sellers into different types, such as "high-frequency high-quality creators," "newcomers," or "sellers whose generation effects need optimization," and match a set of preset interface optimization goals for each type. Subsequently, when the front-end renders the seller backend management page, it can request a style service and obtain a real-time style configuration description file for that seller. Finally, the front-end components dynamically adjust the visual presentation and interaction logic of the control based on this file. For example, for "high-frequency, high-quality creators," controls might be placed in a more prominent position, the interface would be simpler, and their most frequently used style presets would be displayed by default to improve operational efficiency. For "newcomers," controls would strengthen guidance prompts, recommend "intelligent one-click generation" by default, and simplify advanced options to lower the barrier to entry. For "sellers whose performance needs optimization," optimization suggestions could be proactively displayed when they might select inefficient parameters, and high-potential styles would be highlighted. Through this data-driven dynamic adaptation, interactive interfaces best suited to the capabilities and needs of different types of sellers can be provided, thereby effectively improving the usage rate of video generation functions and the quality of the final output content.
[0100] Furthermore, during video generation, a step-by-step wizard-style interactive interface on the seller client 114 can provide a structured guide for sellers through the entire process. This interface employs a state machine-driven front-end architecture; see [link to relevant documentation]. Figures 3 to 9 Initially, a "Video Generation" entry control is displayed on the seller backend management page (e.g., Figure 3(As shown in the "AI Video Creation" control). When a seller triggers this control, a wizard interface is loaded via route redirection or a modal pop-up, clearly displaying all steps of the linear process: ① Fill in / confirm product information, ② Set video attributes (such as style and duration), ③ Select video script, ④ Confirm preview effect and generate. After completing the corresponding operation in each step, the seller can click the "Next" button. At this time, the front-end application will update the current process status through a state management library (such as Vuex or Redux): visually marking the completed steps as completed (such as graying out the font and adding a checkmark icon), while dynamically rendering the input components and instructions for the next step. Intermediate data for each step is temporarily stored on the front end until the last step, when the complete generation parameters are submitted to the task queue of the video platform 112. During the asynchronous video generation by the video platform 112, the seller client 114 can receive the task progress status through WebSocket or a polling mechanism, and update the progress bar and estimated remaining time on the interface in real time; after generation is completed, a notification can be pushed through the message center, and the seller can preview the generated effect online. The entire interaction process significantly reduces the operational complexity for sellers through clear step breakdown, real-time status feedback, and interruption recovery mechanisms, while ensuring the efficient execution and traceability of generated tasks.
[0101] Figure 14 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 14 As shown, device 400 mainly consists of a communication interface 402, a user interface 404, a processor 406, and a data storage 408. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 410. The communication interface 402 enables device 400 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 402 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 402 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 402 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 402 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces.
[0102] User interface 404 includes receiving user input and providing output to the user. Therefore, user interface 404 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 404 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 404 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 400 may support remote access from other devices via communication interface 402 or another physical interface (not shown). User interface 404 may be configured to receive user input, the position and movement of which may be indicated by an indicator or cursor described herein. User interface 404 may also be configured as a display device for rendering or displaying text fragments.
[0103] Processor 406 may contain one or more general-purpose processors and / or special-purpose processors.
[0104] Data storage 408 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 406. Data storage 408 may include removable and non-removable components.
[0105] Processor 406 is capable of executing program instructions 418 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 408 to perform the various functions described herein. Data storage 408 may comprise a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 400, enable device 400 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Processor 406 executing program instructions 418 may result in processor 406 using data 412.
[0106] For example, program instructions 418 may include an operating system 422 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 400 and one or more applications 420 (e.g., a browser, social application, or game application). Similarly, data 412 may include operating system data 416 and application data 414. Operating system data 416 is primarily accessible to the operating system 422, while application data 414 is primarily accessible to one or more applications 420. Application data 414 may reside in a file system visible or hidden from the user of device 400.
[0107] Application 420 can communicate with operating system 422 through one or more application programming interfaces (APIs). These APIs help application 420 read and / or write application data 414, transmit or receive information via communication interface 402, receive or display information on user interface 404, etc.
[0108] In some terminology, application 420 may be simply referred to as "app". Furthermore, application 420 can be downloaded to device 400 through one or more online app stores or app markets. However, applications can also be installed on device 400 in other ways, such as through a web browser or a physical interface on device 400 (e.g., a USB port).
[0109] Please refer to Figure 15 Video generation devices can be applied to, for example Figure 14 The device shown implements the technical solution of this specification. Specifically, the video generation apparatus is applied to a video middleware platform connected to a data server of an e-commerce platform and a seller client of the e-commerce platform. The data server stores product metadata and product materials related to the products sold on the e-commerce platform. The video generation apparatus may include: The receiving module 502 is used to receive a video generation request sent by the seller client; the video generation request carries seller identification information, which is used to indicate the target seller associated with the seller client; The first acquisition module 504 is used to respond to the video generation request, obtain the list of products published by the target seller on the e-commerce platform from the data server based on the seller identification information, and return the product list to the seller client; The second acquisition module 506 is used to, in response to the seller client selecting a target product from the product list, acquire the product metadata and product material set of the target product from the data server, and determine the target product material for generating the video from the product material set based on the product metadata of the target product; The input module 508 is used to input the target product material into the video generation model, so that the video generation model generates a target video based on the target product material.
[0110] Please refer to Figure 16 Video generation devices can be applied to, for example Figure 14The device shown implements the technical solution of this specification. Specifically, the video generation apparatus is applied to a seller client of an e-commerce platform. The seller client is communicatively connected to a video middleware platform, which is also communicatively connected to the data server of the e-commerce platform. The data server stores product metadata and product materials related to the products sold on the e-commerce platform. The video generation apparatus may include: Display module 602 is used to display video generation controls on the seller backend management page; The first sending module 604 is configured to send a video generation request to the video platform in response to detecting an operation command for the video generation control; the video generation request carries seller identification information. The third acquisition module 606 is used to acquire the product list returned by the video middle platform in response to the video generation request, wherein the product list is the list of products published by the target seller to which the seller identification information belongs on the e-commerce platform; The second sending module 608 is used to send the identification information of the target product to the video platform in response to detecting a selection operation for a target product in the product list. This allows the video platform to obtain the product metadata and product material set of the target product from the data server, determine the target product material for generating the video from the product material set based on the product metadata, and input the target product material into the video generation model so that the video generation model generates the target video based on the target product material.
[0111] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0112] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.
[0113] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0114] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0115] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.
[0116] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.
[0117] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0118] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0119] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0120] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0121] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.
Claims
1. A video generation method, applied to a video middleware platform connected to a data server of an e-commerce platform and a seller client of the e-commerce platform, wherein the data server is used to store product metadata and product materials related to products sold on the e-commerce platform; the method includes: Receive the video generation request sent by the seller's client; In response to the video generation request, the system obtains the product metadata and product material set of the target product from the data server, and determines the target product material for generating the video from the product material set based on the product metadata of the target product. The target product material is input into the video generation model, so that the video generation model generates a target video based on the target product material.
2. The method according to claim 1, wherein determining the target product material for generating the video from the product material set based on the product metadata of the target product includes: Retrieve pre-generated thought chain prompts; The thought chain prompt information is used to define a multi-level reasoning step for filtering product materials based on product metadata, and the output of each level of reasoning step serves as the input of at least one level of subsequent reasoning step. The target reasoning model is composed of the thought chain prompts, the product metadata of the target product, and the product material set. The target reasoning model performs thought chain reasoning on the product metadata of the target product and the product material set according to the multi-level reasoning steps, and determines the target product material based on the reasoning results.
3. The method according to claim 2, wherein the multi-level reasoning step comprises: Obtain the task background information for the video generation task. The task background information includes the selection criteria for product materials, the parsing process of product materials, the first evaluation criteria for evaluating whether product materials match product metadata, the user information of the video audience, and the second evaluation criteria for evaluating whether product materials match the video audience. Select multiple candidate product materials that meet the selection criteria from the product material set; The analysis process is described above. Each candidate product material is analyzed to determine the material information of the candidate product material. The first matching degree between the material information of the candidate product material and the product metadata of the target product is evaluated based on the first evaluation criteria. The second matching degree between the material information of the candidate product materials and the user information of the video audience is evaluated based on the second evaluation criteria. The weights of the candidate product materials are determined based on the first matching degree and the second matching degree; Based on the weights of the multiple candidate product materials, the target product material is selected from the multiple candidate product materials.
4. The method according to claim 2, further comprising: Acquire sample data, which includes sample thinking chain prompts, sample product metadata, sample product materials, and tag information. The sample thinking chain prompts include multi-level reasoning steps and the labeled output results corresponding to each level of reasoning. The tag information is used to indicate whether the sample product materials are selected as materials for generating videos. The sample data is input into the base model, so that the base model performs mind chain reasoning on the sample product metadata and the sample product material based on the multi-level reasoning steps in the sample mind chain prompt information, and obtains the reasoning results corresponding to each level of reasoning step; The loss corresponding to each inference step is determined based on the difference between the inference results corresponding to each level of inference steps output by the base model and the labeled output results corresponding to each level of inference steps. The total loss is determined based on the loss corresponding to each inference step. The base model is trained based on the total loss to obtain the target inference model.
5. The method according to claim 1, wherein the video generation request carries seller identification information, the seller identification information being used to represent the target seller associated with the seller client; In response to the video generation request, obtaining the target product's metadata and product material set from the data server includes: Based on the seller identification information, the data server is used to obtain the list of products that the target seller has published on the e-commerce platform, and the list of products is returned to the seller client. In response to the seller client selecting a target product from the product list, the seller client retrieves the target product's metadata and product material set from the data server.
6. The method according to claim 1, after determining the target product material for generating the video from the product material set based on the product metadata of the target product, the method further includes: The product material set and the filtering results of the target product material are returned to the seller's client for display; In response to the seller client's instruction to adjust the target product material, the target product material is adjusted.
7. The method according to claim 1, wherein inputting the target product material into the video generation model, so that the video generation model generates a video based on the target product material, comprises: A video summary is generated based on the product metadata of the target product; The video summary and the target product material are input into the video generation model so that the video generation model generates a target video based on the video summary and the target product material.
8. The method according to claim 7, wherein before inputting the video summary and the target product material into the video generation model, the method further comprises: Return the video summary to the seller's client; In response to receiving the modification instruction from the seller's client, the video summary is modified.
9. The method according to claim 1, wherein inputting the target product material into a video generation model to enable the video generation model to generate a target video based on the target product material comprises: The target product material is input into the video generation model, so that the video generation model generates multiple video preview files based on the target product material; Different video preview files have different visual characteristics; The multiple video preview files are sent to the seller's client for display. In response to receiving a request from the seller client to select a target video preview file from the plurality of video preview files, a target video with the visual features corresponding to the target video preview file is generated.
10. The method according to claim 1, wherein inputting the target product material into a video generation model to generate a target video based on the target product material comprises: Obtain the video attribute information sent by the seller's client; The video attribute information and the target product material are input into the video generation model so that the video generation model generates a target video with the video attribute information based on the target product material.
11. The method according to claim 1, wherein the video platform supports multiple video generation modes, different video generation modes are used to generate different types of target videos, and different types of target videos are generated by different video generation models; The step of inputting the target product material into the video generation model, so that the video generation model generates a target video based on the target product material, includes: Obtain the target video generation mode selected by the seller client from the plurality of video generation modes; The target product material is input into the video generation model corresponding to the target video generation mode, so that the video generation model corresponding to the target video generation mode generates a target video based on the target product material.
12. The method according to claim 1, further comprising: Translate the target video; and / or Perform video splitting on the target video; and / or Publish the target video to the designated platform; and / or Perform video transcoding on the target video; and / or Risk detection is performed on the target video.
13. A video generation method applied to a seller client of an e-commerce platform, wherein the seller client is communicatively connected to a video middleware platform, and the video middleware platform is also communicatively connected to a data server of the e-commerce platform, the data server being used to store product metadata and product materials related to products sold on the e-commerce platform; the method includes: Display the video generation control on the seller backend management page; In response to detecting an operation command for the video generation control, a video generation request is sent to the video platform; Obtain the target video returned by the video platform in response to the video generation request; The target video is generated by the video generation model called by the video platform based on the target product material. The target product material is determined from the product material set of the target product based on the product metadata of the target product. The product metadata and product material set of the target product are obtained from the data server of the e-commerce platform.
14. A video middleware platform, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-12 by executing the executable instructions.
15. A video generation system, comprising an e-commerce platform, a seller client, and a video generation platform; The e-commerce platform includes a data server for storing product metadata and product materials related to the products sold on the e-commerce platform; The seller client is used to send a video generation request to the video platform; The video middleware platform is used to perform the method as described in any one of claims 1-12.
16. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-13.
17. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-13.