Video processing method and device, electronic equipment and storage medium
By using an artificial intelligence model to generate multiple images and descriptive text for items, the video generation process is automated, solving the problem of tedious video generation and improving efficiency and item conversion rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2026-02-14
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies require manual planning, shooting, and editing of materials when generating videos, which is cumbersome and affects the efficiency of video generation.
Using an artificial intelligence model, multiple item images and descriptive text are generated based on a single reference image and detailed information of the item, automatically generating videos that include displays from different angles and scenes.
It eliminates the need for manual shooting and video editing, significantly improving video generation efficiency and publishing speed, and increasing product conversion rates.
Smart Images

Figure CN122120568A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of multimedia technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of multimedia technology, video has become an important tool for describing products. Users increasingly prefer to learn about product features and usage methods through video content, while product providers effectively introduce their products in this format. However, the process of creating videos to introduce products is cumbersome, involving manual planning of video content, shooting video footage, and processing of the footage to generate the final video, which impacts video production efficiency. Summary of the Invention
[0003] This disclosure provides a video processing method, apparatus, electronic device, and storage medium, which not only can display the characteristics of an object in detail, but also solves the problem of insufficient material for the object. It eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation, improving video generation and publishing efficiency, and thus increasing the conversion rate of the object. The technical solution of this disclosure is as follows: According to one aspect of the embodiments of this disclosure, a video processing method is provided, comprising: In response to a video generation instruction for an item, a video generated based on multiple item images and item description text of the item is displayed in the video interface. The multiple item images and the item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to showcase the item from different angles or scenes, and the item description text is used to describe the characteristics of the item. In response to a publishing instruction for the video, the video is published.
[0004] According to one aspect of the embodiments of this disclosure, a video processing method is provided, comprising: In response to a video generation instruction for an item, image prompts and text prompts are generated based on a single reference image and details of the item. The image prompts are used to indicate the content of the item image in the video, and the text prompts are used to indicate the content of the item description text in the video. Based on the image prompts and text prompts, an artificial intelligence model is used to process the reference image and the details information to obtain multiple item images and item description text for the item. The multiple item images are used to display the item from different angles or scenes. Based on the multiple item images and the item description text, a video for showcasing the items is generated.
[0005] According to another aspect of the embodiments of this disclosure, a video processing apparatus is provided, comprising: The display unit is configured to execute a video generation instruction in response to an item, and in the video interface, display a video generated based on multiple item images and item description text of the item. The multiple item images and the item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to display the item from different angles or scenes, and the item description text is used to describe the characteristics of the item. The publishing unit is configured to publish the video in response to a publishing instruction for the video.
[0006] In some embodiments, the display unit is configured to perform at least one of the following: In response to the video generation operation for the item in the video interface, a video generation instruction is generated, and in response to the video generation instruction, a video generated based on multiple item images and item description text of the item is displayed; If the item meets the conditions for generating item content, a video generation instruction is generated, and in response to the video generation instruction, a video generated based on multiple item images and item description text of the item is displayed.
[0007] In some embodiments, the conditions for generating the content of the article include at least one of the following: The inventory quantity of the items is not lower than the quantity threshold; The item is a hot item at the current time, and the hot item is used to indicate the item that is generally needed at the current time. The item has a target identifier, which is used to indicate to an object that possesses the item that needs to distribute the item. The number of searches for keywords related to the item has reached the threshold.
[0008] In some embodiments, the video includes a first split-screen, a second split-screen, and a third split-screen. The first split-screen includes the plurality of object images, the second split-screen includes a close-up image of the object among the plurality of object images, and the third split-screen includes an application image of the object among the plurality of object images. The close-up image is used to display the features of the object, and the application image is used to display the object in an application scenario. The display unit is also configured to play a video of the item in the video interface; during the playback of the video, the first split-screen, the second split-screen, and the third split-screen are displayed sequentially.
[0009] In some embodiments, there are multiple second-shot frames in the video, and different second-shot frames are used to display different close-up images, and different close-up images are used to display the features of the item from different angles; The display unit is further configured to sequentially display multiple second-shot frames during the playback of the video, wherein the display size of the close-up image in each second-shot frame is larger than the display size of the object images in the first-shot frame.
[0010] In some embodiments, the display unit is further configured to display the item description text in the first split-screen during the playback of the video, and to turn off the display of the item description text in the second split-screen and the third split-screen.
[0011] In some embodiments, the display unit is further configured to display the item description text in the form of a sticker during the playback of the video.
[0012] In some embodiments, the plurality of item images include a digital human, which is used to assist in displaying the features of the item, and the digital human has a consistent appearance in each item image in the video.
[0013] In some embodiments, the display unit is further configured to display an item status image of the item when the playback state of the video meets preset conditions, the item status image being used to display real-time information about the item.
[0014] In some embodiments, the display unit is further configured to execute the display of a video content template used to generate the video, the video content template being used to indicate the narrative structure of the video in displaying the item, and the video content template matching the category of the item.
[0015] According to another aspect of the embodiments of this disclosure, a video processing apparatus is provided, comprising: The first generation unit is configured to execute a video generation instruction in response to an item, generating image prompts and text prompts based on a single reference image and details of the item, wherein the image prompts are used to indicate the content of the item image in the video, and the text prompts are used to indicate the content of the item description text in the video; The first processing unit is configured to perform processing on the reference image and the details information based on the image prompts and the text prompts using an artificial intelligence model to obtain multiple item images and item description text of the item, wherein the multiple item images are used to display the item from different angles or scenes; The second generation unit is configured to generate a video for displaying the items based on the plurality of item images and the item description text.
[0016] In some embodiments, the first generation unit is configured to perform the following actions: determining the category of the item based on a reference image of the item and the detailed information; and generating the image prompt and the text prompt based on the category of the item, the reference image, and the detailed information.
[0017] In some embodiments, the first generation unit is configured to perform the following actions: determining a video content template matching the item based on the item's category, the video content template indicating the narrative structure of the video showcasing the item; generating a video script based on the video content template, the reference image, and the details, the video script indicating the content of the video to be generated; and generating the image prompts and text prompts based on the video script.
[0018] In some embodiments, the first generation unit is further configured to generate character prompts if the video content template is a narrative template with a digital human, wherein the digital human is used to assist in displaying the features of the item, and the character prompts are used to indicate that the image of the digital human in each item image in the video remains consistent. The first processing unit is configured to process the reference image based on the image prompt and the person prompt using an artificial intelligence model to obtain multiple item images of the item, wherein the multiple item images include the digital person with a consistent appearance.
[0019] In some embodiments, the apparatus further includes: The second processing unit is configured to perform super-resolution processing on the reference image, extract the main body of the item from the super-resolution processed reference image, and obtain the target image. The first processing unit is configured to perform processing on the target image and the details information based on the image prompts and the text prompts using an artificial intelligence model to obtain multiple item images and item description texts for the item.
[0020] In some embodiments, the first processing unit is configured to, when the length of the image prompt reaches a length threshold, generate an item base map based on the image prompt and the artificial intelligence model, migrate the item in the reference image to the item base map, and obtain an item image of the item, wherein the item base map is used to indicate the scene in which the item is located.
[0021] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the aforementioned video processing method.
[0022] According to another aspect of the present disclosure, a computer-readable storage medium is provided that, when program code in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to perform the video processing method described above.
[0023] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the video processing method described above.
[0024] This disclosure provides a video processing method that, in the process of generating a video for displaying an item, can employ an artificial intelligence model to generate multiple item images and item description texts from different angles or in different scenes based on a single reference image and detailed information of the item. The video is then generated based on these multiple item images and description texts. This not only allows the video to display the characteristics of the item in detail but also solves the problem of insufficient material for the item. Furthermore, it eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation, improving video generation efficiency, and enabling faster video release, thus facilitating higher conversion rates for the item.
[0025] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0027] Figure 1 This is a schematic diagram illustrating the implementation environment of a video processing method according to an exemplary embodiment.
[0028] Figure 2 This is a flowchart illustrating a video processing method according to an exemplary embodiment.
[0029] Figure 3 This is a flowchart illustrating another video processing method according to an exemplary embodiment.
[0030] Figure 4This is a flowchart illustrating another video processing method according to an exemplary embodiment.
[0031] Figure 5 This is a process for generating text stickers according to an exemplary embodiment.
[0032] Figure 6 This is a schematic diagram illustrating a video playback according to an exemplary embodiment.
[0033] Figure 7 This is a schematic diagram illustrating an image of an item's state according to an exemplary embodiment.
[0034] Figure 8 This is a block diagram illustrating a video processing apparatus according to an exemplary embodiment.
[0035] Figure 9 This is a block diagram illustrating another video processing apparatus according to an exemplary embodiment.
[0036] Figure 10 This is a block diagram illustrating a terminal according to an exemplary embodiment.
[0037] Figure 11 This is a block diagram illustrating a server according to an exemplary embodiment. Detailed Implementation
[0038] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0039] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0040] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the reference images and detailed information of the items involved in this disclosure were obtained with full authorization.
[0041] Figure 1 This is a schematic diagram illustrating an implementation environment of a video processing method according to an exemplary embodiment. Taking an electronic device as an example, see [example description missing]. Figure 1 The implementation environment specifically includes: terminal 101 and server 102.
[0042] In some embodiments, terminal 101 is at least one of devices such as a smartphone, smartwatch, desktop computer, laptop, MP3 player, MP4 player, and laptop computer. Terminal 101 has an application that supports video playback installed and running. This application can be a multimedia application, editing application, shopping application, or smart assistant, etc., and this disclosure does not limit this. Users can log in to the application through terminal 101 to access the services provided by the application. Terminal 101 can connect to server 102 via a wireless or wired network, and can then send reference images and detailed information of an item to server 102. Server 102, based on the reference images and detailed information of the item, generates a video for displaying the item using an artificial intelligence model. Terminal 101 can then display the video.
[0043] Terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be several terminals, or dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.
[0044] In some embodiments, server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. Server 102 is used to provide background services for applications that support video playback. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.
[0045] Figure 2 This is a flowchart illustrating a video processing method according to an exemplary embodiment, such as... Figure 2 As shown, the video processing method applied to the terminal includes the following steps: In step 201, in response to the video generation instruction for the item, the terminal displays a video generated based on multiple item images and item description text in the video interface. The multiple item images and item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to display the item from different angles or scenes, and the item description text is used to describe the characteristics of the item.
[0046] In this embodiment, the item can be any item such as clothing, food, or furniture; this embodiment does not limit the scope of the application. Item details may include specifications such as size and color, as well as information such as inventory quantity and price; this embodiment does not limit the scope of the application. Upon detecting a video generation instruction for an item, the server can obtain a single reference image and details of the item, process these using a large language model to generate multiple item images and item description text, and then generate a video to display the item based on these multiple images and description text. Finally, the terminal displays the generated video in the video interface to demonstrate the video generation effect to the user.
[0047] The video generation instruction can be generated when a video generation operation is detected in the video interface, or it can be generated when an item meets a certain condition. That is, the video processing method provided in this solution can be initiated by the user or automatically executed when an item meets a certain condition; this disclosure does not limit this. Multiple item images are generated by extending a single reference image of the item, showcasing the characteristics of the item from different angles or scenes (such as a specific application scenario). The item description text is generated from the item's detailed information.
[0048] In step 202, in response to a publishing instruction for the video, the terminal publishes the video.
[0049] In this embodiment of the disclosure, the terminal publishes a video upon detecting a video publishing command. The video interface may display a video publishing control; when the video publishing control is triggered, the terminal publishes the video.
[0050] This disclosure provides a video processing method that, in the process of generating a video for displaying an item, can employ an artificial intelligence model to generate multiple item images and item description texts from different angles or in different scenes based on a single reference image and detailed information of the item. The video is then generated based on these multiple item images and description texts. This not only allows the video to display the characteristics of the item in detail but also solves the problem of insufficient material for the item. Furthermore, it eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation, improving video generation efficiency, and enabling faster video release, thus facilitating higher conversion rates for the item.
[0051] In some embodiments, in response to a video generation instruction for an item, a video generated based on multiple item images and item description text is displayed in the video interface, including at least one of the following: In response to the video generation operation for an item in the video interface, a video generation command is generated, and in response to the video generation command, a video generated based on multiple item images and item description text is displayed; If an item meets the conditions for generating item content, a video generation instruction is generated. In response to the video generation instruction, a video generated based on multiple item images and item description text is displayed.
[0052] The solution provided in this disclosure offers two triggering methods for video generation instructions: it supports users actively triggering video generation operations on the video interface, and it also supports the system automatically triggering when an item meets preset content generation conditions. This dual-mode triggering mechanism makes the video generation process highly adaptable and flexible, allowing the selection of the most efficient startup method according to different application scenarios. This significantly reduces manual operation costs and achieves full automation from condition judgment to content production, thereby significantly improving video generation efficiency and ensuring that important items can receive timely video display support.
[0053] In some embodiments, the conditions for generating the content of an item include at least one of the following: The inventory quantity of the items is not lower than the quantity threshold; The item is a hot item at the current time. Hot items are used to indicate items that are generally needed at the current time. The item has a target identifier, which indicates that the object possessing the item needs to be distributed to it. The number of searches for keywords related to the item has reached the threshold.
[0054] The solution provided in this disclosure clarifies multiple specific dimensions for judging the conditions for generating content for items, including inventory quantity thresholds, current hot items, the target identifier status of items, and the search popularity of keywords related to items. These conditions form an intelligent video generation decision system, enabling the system to autonomously determine which items need to be prioritized for video generation based on multi-dimensional data analysis. This avoids the subjectivity and lag of manual judgment, greatly reduces the time consumed by manual screening and evaluation, and significantly improves the efficiency of target selection for video generation. It ensures that limited computing and attention resources are invested in the most valuable items, thereby optimizing the overall video generation process.
[0055] In some embodiments, the video includes a first split-screen, a second split-screen, and a third split-screen. The first split-screen includes multiple object images, the second split-screen includes close-up images of the objects among the multiple object images, and the third split-screen includes application images of the objects among the multiple object images. The close-up images are used to show the features of the objects, and the application images are used to show the objects in application scenarios. The method also includes: In the video interface, play the video of the item; During the video playback, the first, second, and third storyboard scenes are displayed sequentially.
[0056] The solution provided in this disclosure defines a standardized storyboard structure for videos, including a first storyboard containing multiple object images, a second storyboard featuring close-up images highlighting the object's characteristics, and a third storyboard showcasing the application scenario. It also specifies their ordered playback sequence. This structured narrative framework provides a clear logical organization for video content, allowing content with different display purposes to be rationally allocated along the timeline. Furthermore, by pre-setting the storyboard type and playback order, it eliminates the need to redesign the content structure each time the video is generated. Materials can be directly organized according to a templated narrative logic, greatly simplifying the complexity of video generation and reducing the time spent on video content planning and arrangement, thereby significantly improving the overall efficiency of video generation.
[0057] In some embodiments, there are multiple second-shot frames in the video, and different second-shot frames are used to display different close-up images, which are used to display the features of the object from different angles. The method also includes: During the video playback, multiple second-scene frames are displayed sequentially, with the display size of the close-up images in each second-scene frame being larger than the display size of the images of each item in the first-scene frame.
[0058] The solution provided in this disclosure refines the implementation scheme for displaying close-up images. By setting multiple second-shot frames, each frame displays close-up images taken from different angles, and the display size of the close-up images is larger than the object images in the first-shot frame (panoramic view). This multi-angle, large-format close-up display method can systematically present the detailed features of the object, avoiding the problem of missing important information in a single close-up. Moreover, without manual intervention, this automated multi-angle close-up generation mechanism greatly simplifies the planning work for displaying the details of the object, significantly improves the completeness and systematicness of feature display, and thus improves video generation efficiency.
[0059] In some embodiments, the method further includes: During video playback, the item description text is displayed in the first shot, and is turned off in the second and third shots.
[0060] The solution provided in this disclosure specifies that the text description of the item is displayed in the first storyboard, while the text display is turned off in the second and third storyboards. This differentiated text display scheme allows the text information and visual content to be optimally matched according to the purpose of the storyboard. That is, the text is displayed in the panoramic view (first storyboard) to provide an overall introduction, while the text is avoided from interfering with the visual focus in close-up and scene shots, which facilitates better display of the characteristics of the item and improves the video quality.
[0061] In some embodiments, the method further includes: Item description text is displayed as stickers during video playback.
[0062] The solution provided in this disclosure displays item description text in the form of stickers during video playback. This not only has stronger visual appeal, making it easier for viewers to discover the characteristics of the items and improving information transmission efficiency, but also allows the item description text to be better integrated into the video frame without being too abrupt, improving the aesthetics and consistency of the text elements, reducing the workload of designing text visual effects, and thus improving video generation efficiency and video quality.
[0063] In some embodiments, multiple object images include a digital human, which is used to help demonstrate the features of the objects, and the digital human appears consistently in all object images in the video.
[0064] The solution provided in this disclosure introduces a digital human as an auxiliary element for displaying items, and ensures that the image of the digital human remains consistent across various item images in the video. This digital human-assisted display method can add anthropomorphic elements to the item display, making the characteristics of the items more vividly presented through the interactive behavior of the digital human. By ensuring the consistency of the digital human image across scenes, a continuous visual narrative effect can be created, which can significantly improve the coherence of video content and video quality.
[0065] In some embodiments, the method further includes: When the video playback status meets the preset conditions, the item status image is displayed, which is used to show the real-time information of the item.
[0066] The solution provided in this disclosure displays an image reflecting the real-time information of an item when the playback status meets preset conditions. This allows the video to not only showcase the item itself but also reflect its availability in real time, providing viewers with more comprehensive decision-making information and improving information transmission efficiency. Furthermore, by associating the supply status information with the video playback status, it can intelligently determine when to insert status information, avoiding information overload or inappropriate display timing. This automated status information integration mechanism can significantly improve the dimensionality and real-time performance of video information, reduce the workload of manually adding and updating status information, and thus improve the efficiency of the dynamic information integration process in video production.
[0067] In some embodiments, the method further includes: This displays the video content template used to generate the video. The video content template is used to indicate the narrative structure of the items shown in the video, and the video content template matches the category of the item.
[0068] The solution provided in this disclosure presents a video content template to the user during the video generation process. This template indicates the narrative structure of the items displayed in the video and matches the item categories, enabling the user to understand the structured framework of the video generation. This helps the user quickly understand the video's organizational logic and improves information transmission efficiency.
[0069] The above Figure 2 The above only shows the video processing method on the terminal side. The following section introduces the video processing method provided in this disclosure from the server side. Figure 3 This is a flowchart illustrating another video processing method according to an exemplary embodiment, such as... Figure 3 As shown, the video processing method applied to the server includes the following steps: In step 301, in response to a video generation instruction for an item, the server generates image cues and text cues based on a single reference image and details of the item. The image cues indicate the content of the item image in the video, and the text cues indicate the content of the item description text in the video.
[0070] In this embodiment of the disclosure, upon receiving a video generation instruction for an item, the server obtains a single reference image and detailed information of the item, and generates image prompts and text prompts based on the reference image and detailed information, so as to subsequently instruct the artificial intelligence model to generate the content of the required item image and the content of the item description text in the video.
[0071] In step 302, the server processes the reference image and details based on the image prompts and text prompts using an artificial intelligence model to obtain multiple item images and item description text. The multiple item images are used to display the item from different angles or scenes.
[0072] In this embodiment of the disclosure, the server inputs image prompts, text prompts, reference images, and detailed information into the artificial intelligence model. The image prompts and text prompts instruct the artificial intelligence model to process the reference images and detailed information, thereby generating multiple item images and item description text for an item.
[0073] In step 303, the server generates a video to showcase the items based on multiple item images and item description text.
[0074] In this embodiment, the server can use an intelligent orchestration algorithm to arrange item images chronologically according to a preset narrative logic, forming the visual main body of the video; it can parse the item description text, extract key information points, and automatically generate matching voice narration or subtitle content. Subsequently, the server can intelligently match background music, transition effects, and dynamic text stickers for the video to ensure a smooth and consistent audiovisual experience. Finally, the server performs multi-track synthesis and encoding compression of the audio track, image sequence, and special effects elements, outputting a standard format video file that can be directly distributed.
[0075] This disclosure provides a video processing method. In the process of generating a video for displaying an item, image prompts indicating the content of the item image and text prompts indicating the content of the item description text are first generated. Then, these prompts guide an artificial intelligence model to process reference images and detailed information. This prompt-driven generation method transforms the content creation intention into precise instructions that the machine can understand, achieving precise control over the generation process. By generating prompts step by step before content generation, the complex creative task is broken down into more controllable sub-tasks, improving the predictability and stability of the generation process. This significantly improves video quality, solves the problem of insufficient material for the item, and eliminates the need for manual shooting and editing of materials, effectively lowering the threshold for video generation and improving video generation efficiency. This allows for faster video release and improves the conversion rate of the item.
[0076] In some embodiments, image prompts and text prompts are generated based on a single reference image and details of the item, including: Based on the reference image and detailed information of the item, determine the item's category; Based on the item's category, reference image, and details, generate image and text prompts.
[0077] The solution provided in this disclosure introduces an item category judgment step in the prompt word generation process. Based on the item category, reference image, and details, image prompt words and text prompt words are generated. This category-aware prompt word generation strategy enables more targeted creation instructions to be generated according to the characteristics of different item categories, avoiding the problem of poor video generation effect that may be caused by "one-size-fits-all" general instructions, and improving video quality.
[0078] In some embodiments, image prompts and text prompts are generated based on the item's category, reference image, and details, including: Based on the category of the item, determine the video content template that matches the item. The video content template is used to indicate the narrative structure of the video showcasing the item. Based on the video content template, reference images, and detailed information, a video script is generated. The video script is used to indicate the content of the video to be generated. Generate image and text prompts based on the video script.
[0079] The solution provided in this disclosure constructs a multi-layered content planning system, from video content templates to video scripts to prompts. From macro-narrative structure to specific content planning to machine execution instructions, it can ensure the quality of generated content in terms of narrative logic, content integrity, and execution feasibility. It avoids the structural chaos that may result from skipping directly from raw materials to finished product, significantly improves the organization and integrity of video content, reduces the time consumption of content structure design, and thus improves video generation efficiency and ensures video quality.
[0080] In some embodiments, the method further includes: If the video content template is a narrative template with digital humans, then character prompts are generated. The digital humans are used to help display the characteristics of the items, and the character prompts are used to indicate that the image of the digital humans in the various item images in the video remains consistent. The process of obtaining multiple images of an item by processing a reference image based on image prompts and using an artificial intelligence model includes: Based on image and person cues, an artificial intelligence model is used to process reference images to obtain multiple object images, each containing a digitally identical person.
[0081] The solution provided in this disclosure allows the system to generate specific character prompts when a video content template requires digital human assistance for display. These prompts, along with image prompts, guide the model to generate object images containing a consistent digital human image. By setting up an independent prompt control channel for the digital human, the system can more accurately manage the digital human's image characteristics and prevent changes in image in complex scenes. This specialized character image control mechanism can significantly improve the professionalism and credibility of digital human-assisted object display, reduce the technical difficulty of maintaining consistency in digital human image across multiple scenarios, and thus improve video generation efficiency and video quality.
[0082] In some embodiments, before generating multiple item images based on a reference image, the method further includes: Super-resolution processing is performed on the reference image, and the main body of the object is extracted from the super-resolution processed reference image to obtain the target image; Based on image and text prompts, an artificial intelligence model processes reference images and detailed information to obtain multiple item images and item description texts, including: Based on image and text prompts, an artificial intelligence model is used to process the target image and details to obtain multiple item images and item description text.
[0083] The solution provided in this disclosure adds an image preprocessing step before generating an object image based on a reference image. This includes super-resolution processing of the reference image and object subject extraction. Super-resolution processing improves image clarity, while object extraction yields a clean object image without background. This ensures that the subsequent generation process is based on a high-quality visual foundation, avoiding the problem of reduced generation effect due to low quality of the original image, and can significantly improve video quality.
[0084] In some embodiments, the process of obtaining an item image by processing a reference image based on image prompts using an artificial intelligence model includes: When the length of the image prompt reaches a length threshold, an item base map is generated based on the image prompt and the artificial intelligence model. The item in the reference image is then transferred to the item base map to obtain the item image. The item base map is used to indicate the scene in which the item is located.
[0085] The solution provided in this disclosure proposes a two-stage processing strategy for complex image generation tasks. When the image prompt is too long, a scene base map is first generated based on the image prompt, and then the items in the reference image are migrated to the base map. This approach decomposes the complex image generation task into two relatively independent sub-tasks: scene generation and item migration. This avoids the problem of decreased model compliance that may be caused by processing too many requirements at the same time in a single generation, and can significantly improve the success rate and quality stability of complex item image generation, thereby improving video quality.
[0086] The above Figure 2 and Figure 3 The diagram shown is merely the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on a specific implementation method. Figure 4 This is a flowchart illustrating another video processing method according to an exemplary embodiment. Taking an electronic device provided as a terminal and a server as an example, see [link to example]. Figure 4 The video processing method includes: In step 401, in response to a video generation instruction for an item, the server generates image cues and text cues based on a single reference image and details of the item. The image cues indicate the content of the item image in the video, and the text cues indicate the content of the item description text in the video.
[0087] In this embodiment of the disclosure, upon receiving a video generation instruction for an item, the server obtains a single reference image and detailed information of the item, and generates an image prompt and a text prompt based on the reference image and detailed information, so as to subsequently instruct the artificial intelligence model to generate the content of the required item image and the content of the item description text in the video.
[0088] In some embodiments, a video generation instruction can be generated when a video generation operation is detected in the video interface, or it can be generated when an item meets a certain condition. Optionally, a video generation instruction is generated in response to a video generation operation for an item in the video interface. Alternatively, a video generation instruction is generated when the item meets the item content generation conditions. The solution provided by the embodiments of this disclosure offers two triggering methods for video generation instructions: it supports both user-initiated video generation operations in the video interface and automatic triggering by the system when an item meets preset item content generation conditions. This dual-mode triggering mechanism makes the video generation process highly adaptable and flexible, allowing the selection of the most efficient startup method according to different application scenarios. This significantly reduces manual operation costs, achieves full-process automation from condition judgment to content production, and can significantly improve video generation efficiency, ensuring that important items can obtain video display support in a timely manner.
[0089] This disclosure does not limit the specific content of the conditions for generating the article content. Optionally, the conditions for generating the article content may include at least one of the following, but are by no means limited thereto.
[0090] The first requirement is that the inventory quantity of the item is not lower than the quantity threshold.
[0091] The second point is that the item is a hot item at the current time. Hot items are used to indicate items that are generally needed at the current time. The current time can be a season, and correspondingly, the item is an item applicable to the current season; or, the current time can be a holiday, and correspondingly, the item is an item that matches the holiday, etc., and this disclosure does not limit this aspect.
[0092] Thirdly, the item has a target identifier. The target identifier is used to indicate to an object that possesses the item that it needs to be distributed. This target identifier can be a new product identifier, a promotional identifier, or a clearance identifier, etc., and this embodiment of the disclosure does not limit it in this way.
[0093] The fourth item is that the number of searches for keywords related to the item reaches the threshold.
[0094] The solution provided in this disclosure clarifies multiple specific dimensions for judging the conditions for generating content for items, including inventory quantity thresholds, current hot items, the target identifier status of items, and the search popularity of keywords related to items. These conditions form an intelligent video generation decision system, enabling the system to autonomously determine which items need to be prioritized for video generation based on multi-dimensional data analysis. This avoids the subjectivity and lag of manual judgment, greatly reduces the time consumed by manual screening and evaluation, and significantly improves the efficiency of target selection for video generation. It ensures that limited computing and attention resources are invested in the most valuable items, thereby optimizing the overall video generation process.
[0095] In the process of generating image and text prompts, the server can determine the category of an item based on its reference image and details. Then, the server generates image and text prompts based on the item's category, reference image, and details. The solution provided in this disclosure introduces an item category determination step in the prompt generation process, generating image and text prompts based on the item's category, reference image, and details. This category-aware prompt generation strategy enables the generation of more targeted creative instructions based on the characteristics of different item categories, avoiding the poor video generation results that may result from a "one-size-fits-all" approach and improving video quality.
[0096] The server can determine a video content template matching the item's category. This template guides the narrative structure of the video showcasing the item. Then, based on the video content template, reference images, and detailed information, the server generates a video script, which indicates the content of the video to be generated. Next, based on the video script, the server generates image and text prompts. The server can store multiple video content templates suitable for showcasing different types of items. For example, a food content template might first show the food through close-up images, then depict its preparation process and environment, followed by descriptions of taste and the user's enjoyment. Similarly, a house content template might first show the house's exterior, then its interior details, and finally the user's stay experience.
[0097] The solution provided in this disclosure constructs a multi-layered content planning system, from video content templates to video scripts to prompts. From macro-narrative structure to specific content planning to machine execution instructions, it can ensure the quality of generated content in terms of narrative logic, content integrity, and execution feasibility. It avoids the structural chaos that may result from skipping directly from raw materials to finished product, significantly improves the organization and integrity of video content, reduces the time consumption of content structure design, and thus improves video generation efficiency and ensures video quality.
[0098] In step 402, the server processes the reference image and details based on the image prompts and text prompts using an artificial intelligence model to obtain multiple item images and item description text. The multiple item images are used to display the item from different angles or scenes.
[0099] In this embodiment of the disclosure, the server inputs image prompts, text prompts, reference images, and detailed information into the artificial intelligence model. The image prompts and text prompts instruct the artificial intelligence model to process the reference images and detailed information, thereby generating multiple item images and item description text for an item.
[0100] In some embodiments, if the video content template is a narrative template with a digital human, the server can generate character prompts. The digital human is used to assist in showcasing the features of the item, and the character prompts are used to indicate that the image of the digital human in each item image in the video remains consistent. Then, during the generation of item images, the server processes the reference image based on the image prompts and character prompts using an artificial intelligence model to obtain multiple item images. These multiple item images include a digital human with a consistent image. The solution provided in this disclosure, when the video content template requires the assistance of a digital human for display, generates specialized character prompts, which, in conjunction with the image prompts, guide the model to generate item images containing a consistent digital human image. By setting up an independent prompt control channel for the digital human, the image features of the digital human can be managed more precisely, avoiding changes in image in complex scenes. This specialized character image control mechanism can significantly improve the professionalism and credibility of digital human-assisted item display, reduce the technical difficulty of maintaining the consistency of the digital human image in multiple scenes, thereby improving video generation efficiency and video quality.
[0101] In some embodiments, before generating multiple item images based on a reference image, the server can perform super-resolution processing on the reference image to extract the main body of the item from the super-resolution processed reference image, obtaining a target image. Then, during the generation of item images, the server processes the target image and details information based on image prompts and text prompts using an artificial intelligence model to obtain multiple item images and item description text. The solution provided by this disclosure adds an image preprocessing step before generating item images based on a reference image, including super-resolution processing of the reference image and extraction of the main body of the item. Super-resolution processing improves image clarity, while main body extraction yields a clean item image with background removed, ensuring that the subsequent generation process is based on a high-quality visual foundation. This avoids the problem of decreased generation effect due to low quality of the original image and can significantly improve video quality.
[0102] In some embodiments, when the length of the image cues reaches a length threshold, the server can generate an object base map based on the image cues and an artificial intelligence model, and then transfer the object from the reference image to the object base map to obtain an object image. The object base map is used to indicate the scene in which the object is located. The solution provided in this disclosure proposes a two-stage processing strategy for complex image generation tasks. When the image cues are too long, a scene base map is first generated based on the image cues, and then the object from the reference image is transferred to the base map. This approach decomposes the complex image generation task into two relatively independent sub-tasks: scene generation and object transfer. This avoids the problem of decreased model compliance that may result from processing too many requirements simultaneously in a single generation, and can significantly improve the success rate and quality stability of complex object image generation, thereby improving video quality.
[0103] In step 403, the server generates a video to showcase the items based on multiple item images and item description text.
[0104] In this embodiment, the server can use an intelligent orchestration algorithm to arrange item images chronologically according to a preset narrative logic, forming the visual main body of the video; it parses the item description text, extracts key information points, and automatically generates matching voice narration or subtitle content. Subsequently, the server can intelligently match background music, transition effects, and dynamic text stickers for the video to ensure a smooth and consistent audiovisual experience. Finally, the server performs multi-track synthesis and encoding compression of the audio track, image sequence, and special effects elements, outputting a standard format video file that can be directly distributed. Specifically, the server can generate text stickers of item description text and then merge these text stickers with multiple item images to generate a video for showcasing the items.
[0105] For example, Figure 5 This describes a process for generating text stickers according to an exemplary embodiment. See also: Figure 5The generation of text stickers begins in the input configuration stage. The server receives specular SVG (Scalable Vector Graphics) vector files, custom font files, and HTML style templates as core materials. Under a multi-layered cross-platform compatible architecture, Playwright (an open-source automation library) drives the Chromium (an open-source browser project) kernel to achieve accurate web-based text rendering. PIL (Python Imaging Library) and Pilmoji (a PIL-based Python library specifically designed for correctly rendering colored emojis in images) collaboratively process the compositing of the basic image and emojis. The engine then parses the font file and decodes the glyph data, combining this with CSS (Cascading Style Sheets) style directives from the HTML template to render a text layer with complex styles in Chromium. Simultaneously, the system calculates the lighting range based on the specular SVG path and generates a corresponding transparency mask layer. Finally, through multi-layer compositing technology, the text layer, specular mask, and emojis are blended pixel-by-pixel, outputting a professional-grade text sticker image with dynamic visual effects and a transparent background.
[0106] In step 404, the server sends video to the terminal.
[0107] In step 405, the terminal displays a video generated based on multiple item images and item description text in the video interface.
[0108] In this embodiment of the disclosure, the terminal displays the generated video in the video interface to demonstrate the video generation effect to the user.
[0109] In some embodiments, the video includes a first split-screen, a second split-screen, and a third split-screen. The first split-screen includes multiple object images, the second split-screen includes close-up images of the objects among the multiple object images, and the third split-screen includes application images of the objects among the multiple object images. The close-up images are used to demonstrate the features of the objects, and the application images are used to demonstrate the objects in an application scenario. Accordingly, the terminal plays the video of the objects in the video interface. Then, during the video playback, the terminal sequentially displays the first split-screen, the second split-screen, and the third split-screen.
[0110] The solution provided in this disclosure defines a standardized storyboard structure for videos, including a first storyboard containing multiple object images, a second storyboard featuring close-up images highlighting the object's characteristics, and a third storyboard showcasing the application scenario. It also specifies their ordered playback sequence. This structured narrative framework provides a clear logical organization for video content, allowing content with different display purposes to be rationally allocated along the timeline. Furthermore, by pre-setting the storyboard type and playback order, it eliminates the need to redesign the content structure each time the video is generated. Materials can be directly organized according to a templated narrative logic, greatly simplifying the complexity of video generation and reducing the time spent on video content planning and arrangement, thereby significantly improving the overall efficiency of video generation.
[0111] The video can have multiple second-shot frames. Different second-shot frames are used to display different close-up images, and different close-up images are used to showcase the features of the object from different angles. Accordingly, during video playback, the terminal sequentially displays multiple second-shot frames, with the display size of the close-up images in each second-shot frame being larger than the display size of the object images in the first-shot frame. The solution provided by this disclosure refines the implementation scheme for displaying close-up images. By setting multiple second-shot frames, each frame displays close-up images taken from different angles, and the display size of the close-up images is larger than the object images in the first-shot frame (panoramic view), this multi-angle, large-format close-up display method can systematically present the detailed features of the object, avoiding the problem of missing important information in a single close-up; and without manual intervention, this automated multi-angle close-up generation mechanism greatly simplifies the planning work for displaying the details of the object, significantly improving the completeness and systematicness of feature display, thereby improving video generation efficiency.
[0112] In some embodiments, during video playback, the terminal displays item description text in the first shot and disables the display of item description text in the second and third shots. The solution provided by this disclosure specifies that item description text is displayed in the first shot, while its display is disabled in the second and third shots. This differentiated text display scheme allows for optimal matching of text information and visual content according to the purpose of each shot. Specifically, text is displayed in panoramic shots (first shot) to provide an overall introduction, while text is avoided in close-ups and scene shots to prevent visual distraction, facilitating better display of item characteristics and improving video quality.
[0113] In some embodiments, the terminal displays item description text as stickers during video playback. The solution provided in this disclosure, which displays item description text as stickers during video playback, not only has stronger visual appeal, making it easier for viewers to quickly identify item features and improving information transmission efficiency, but also allows the item description text to blend better into the video frame without being too abrupt, improving the aesthetics and consistency of text elements, reducing the workload of text visual effects design, thereby improving video generation efficiency and video quality.
[0114] For example, Figure 6 This is a schematic diagram illustrating a video playback according to an exemplary embodiment. See also... Figure 6 The first shot displays four images of an item arranged into a four-square grid, with sticker-style descriptions. This high-information-density presentation quickly conveys all the information to the user, simplifying the decision-making process. The second shot then shows close-up images of the items to highlight their detailed features. Finally, the third shot shows a posed image of the item being held, demonstrating its practical use.
[0115] In some embodiments, multiple object images include a digital human. The digital human is used to assist in showcasing the features of the objects, and the digital human's appearance remains consistent across all object images in the video. The solution provided by this disclosure introduces a digital human as an auxiliary element for object display and ensures that the digital human's appearance remains consistent across all object images in the video. This digital human-assisted display method adds anthropomorphic elements to the object display, making the object features more vividly presented through the interactive behavior of the digital human. By ensuring the consistency of the digital human's image across scenes, a continuous visual narrative effect can be created, significantly improving the coherence of the video content and the video quality.
[0116] In some embodiments, when the video playback status meets preset conditions, the terminal displays an item status image, which is used to display the real-time information of the item. The item status image may display information such as the item's inventory quantity, price, coupons (or vouchers), etc., which are not limited in this embodiment. The preset conditions may be that the video playback duration reaches a preset duration, or that the video plays to a specific item image, etc., which are not limited in this embodiment. The solution provided by this embodiment displays an image reflecting the real-time information of the item when the playback status meets preset conditions, so that the video not only displays the item itself but also reflects the item's availability status in real time, providing viewers with more comprehensive decision-making information and improving information transmission efficiency. Furthermore, by associating the supply status information with the video playback status, it can intelligently determine when to insert status information, avoiding information overload or inappropriate display timing. This automated status information integration mechanism can significantly improve the dimensionality and real-time nature of video information, reduce the workload of manually adding and updating status information, and thus improve the efficiency of the dynamic information integration stage in video production.
[0117] For example, Figure 7 This is a schematic diagram illustrating an image of an article's state according to an exemplary embodiment. See also... Figure 7 The item status image 701 shows the availability of vouchers associated with the item.
[0118] In some embodiments, the terminal can also display a video content template used to generate the video. The video content template is used to indicate the narrative structure of the items displayed in the video, and the video content template matches the category of the items. The solution provided by the embodiments of this disclosure displays the video content template used to the user during the video generation process. This template indicates the narrative structure of the items displayed in the video and matches the category of the items, enabling the user to understand the structured framework of the video generation, thereby helping the user quickly understand the organizational logic of the video and improving the efficiency of information transmission.
[0119] In step 406, in response to a publishing instruction for the video, the terminal publishes the video.
[0120] In this embodiment of the disclosure, the terminal publishes a video upon detecting a video publishing command. The video interface may display a video publishing control; when the video publishing control is triggered, the terminal publishes the video.
[0121] This disclosure provides a video processing method that, in the process of generating a video for displaying an item, can employ an artificial intelligence model to generate multiple item images and item description texts from different angles or in different scenes based on a single reference image and detailed information of the item. The video is then generated based on these multiple item images and description texts. This not only allows the video to display the characteristics of the item in detail but also solves the problem of insufficient material for the item. Furthermore, it eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation, improving video generation efficiency, and enabling faster video release, thus facilitating higher conversion rates for the item.
[0122] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0123] Figure 8 This is a block diagram illustrating a video processing apparatus according to an exemplary embodiment. See also: Figure 8 The video processing device includes: Display unit 801 is configured to execute a video generation instruction in response to an item, and in the video interface, display a video generated based on multiple item images and item description text of the item. The multiple item images and item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to display the item from different angles or scenes, and the item description text is used to describe the characteristics of the item. Publishing unit 802 is configured to publish video in response to a publishing command for video.
[0124] In some embodiments, the display unit 801 is configured to perform at least one of the following: In response to the video generation operation for an item in the video interface, a video generation command is generated, and in response to the video generation command, a video generated based on multiple item images and item description text is displayed; If an item meets the conditions for generating item content, a video generation instruction is generated. In response to the video generation instruction, a video generated based on multiple item images and item description text is displayed.
[0125] In some embodiments, the conditions for generating the content of an item include at least one of the following: The inventory quantity of the items is not lower than the quantity threshold; The item is a hot item at the current time. Hot items are used to indicate items that are generally needed at the current time. The item has a target identifier, which indicates that the object possessing the item needs to be distributed to it. The number of searches for keywords related to the item has reached the threshold.
[0126] In some embodiments, the video includes a first split-screen, a second split-screen, and a third split-screen. The first split-screen includes multiple object images, the second split-screen includes close-up images of the objects among the multiple object images, and the third split-screen includes application images of the objects among the multiple object images. The close-up images are used to show the features of the objects, and the application images are used to show the objects in application scenarios. The display unit 801 is also configured to play a video of an item in a video interface; during the video playback, the first split-screen, the second split-screen, and the third split-screen are displayed sequentially.
[0127] In some embodiments, there are multiple second-shot frames in the video, and different second-shot frames are used to display different close-up images, which are used to display the features of the object from different angles. The display unit 801 is also configured to sequentially display multiple second-shot frames during video playback, wherein the display size of the close-up image in each second-shot frame is larger than the display size of the object images in the first-shot frame.
[0128] In some embodiments, the display unit 801 is further configured to display item description text in a first shot during video playback, and to turn off the display of item description text in a second shot and a third shot.
[0129] In some embodiments, the display unit 801 is also configured to display item description text in the form of stickers during video playback.
[0130] In some embodiments, multiple object images include a digital human, which is used to help demonstrate the features of the objects, and the digital human appears consistently in all object images in the video.
[0131] In some embodiments, the display unit 801 is further configured to display an item status image when the video playback state meets preset conditions, the item status image being used to display real-time information about the item.
[0132] In some embodiments, the display unit 801 is further configured to execute a video content template used to display the generated video, the video content template being used to indicate the narrative structure of the items displayed in the video, and the video content template being matched with the category of the item.
[0133] This disclosure provides a video processing apparatus that, in the process of generating a video for displaying an item, can employ an artificial intelligence model to generate multiple item images and item description text from different angles or in different scenes based on a single reference image and detailed information of the item. The video is then generated based on these multiple item images and description text. This not only allows the video to display the characteristics of the item in detail but also solves the problem of insufficient material for the item. Furthermore, it eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation, improving video generation efficiency, and enabling faster video release, thus facilitating a higher conversion rate for the item.
[0134] Figure 9 This is a block diagram illustrating another video processing apparatus according to an exemplary embodiment. See also Figure 9 The video processing device includes: The first generation unit 901 is configured to execute a video generation instruction for an item in response to the item, and generate image cues and text cues based on a single reference image and detail information of the item. The image cues are used to indicate the content of the item image in the video, and the text cues are used to indicate the content of the item description text in the video. The first processing unit 902 is configured to perform processing based on image prompts and text prompts, using an artificial intelligence model to process reference images and details to obtain multiple item images and item description text. The multiple item images are used to display the item from different angles or scenes. The second generation unit 903 is configured to generate a video for displaying items based on multiple item images and item description text.
[0135] In some embodiments, the first generation unit is configured to perform actions based on the item's reference image and details to determine the item's category; and to generate image prompts and text prompts based on the item's category, reference image, and details.
[0136] In some embodiments, the first generation unit is configured to perform the following: based on the item category, determine a video content template matching the item, the video content template being used to indicate the narrative structure of the video showcasing the item; based on the video content template, reference images, and details, generate a video script, the video script being used to indicate the content of the video to be generated; and based on the video script, generate image prompts and text prompts.
[0137] In some embodiments, the first generation unit is further configured to generate character prompts if the video content template is a narrative template with digital humans, wherein the digital humans are used to assist in displaying the features of the items, and the character prompts are used to indicate that the image of the digital humans in each item image in the video remains consistent. The first processing unit is configured to process reference images based on image prompts and person prompts using an artificial intelligence model to obtain multiple object images, including a digital person with a consistent appearance.
[0138] In some embodiments, the apparatus further includes: The second processing unit is configured to perform super-resolution processing on the reference image, extract the main body of the object from the super-resolution processed reference image, and obtain the target image. The first processing unit is configured to perform processing based on image prompts and text prompts, using an artificial intelligence model to process the target image and details information to obtain multiple item images and item description text.
[0139] In some embodiments, the first processing unit is configured to generate an item base map based on the image cues and an artificial intelligence model when the length of the image cues reaches a length threshold, migrate the item in the reference image to the item base map, and obtain an item image of the item. The item base map is used to indicate the scene in which the item is located.
[0140] This disclosure provides a video processing apparatus that, in the process of generating a video for displaying an item, first generates image prompts indicating the content of the item's image and text prompts indicating the content of the item's descriptive text. These prompts then guide an artificial intelligence model to process reference images and detailed information. This prompt-driven generation method transforms the content creation intent into precise machine-understandable instructions, achieving precise control over the generation process. By generating prompts step-by-step before content generation, the complex creative task is broken down into more controllable sub-tasks, improving the predictability and stability of the generation process. This significantly improves video quality, solves the problem of insufficient material for the item, and eliminates the need for manual shooting and editing of materials, effectively lowering the barrier to video generation and improving video generation efficiency. This allows for faster video release and facilitates higher conversion rates for the item.
[0141] It should be noted that the video processing apparatus provided in the above embodiments is only illustrated by the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the video processing apparatus and video processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0142] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0143] When an electronic device is provided as a terminal, Figure 10 This is a block diagram illustrating a terminal 1000 according to an exemplary embodiment. The terminal... Figure 10 A structural block diagram of a terminal 1000 provided in an exemplary embodiment of this disclosure is shown. The terminal 1000 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1000 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0144] Typically, terminal 1000 includes a processor 1001 and a memory 1002.
[0145] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0146] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one computer program, which is executed by the processor 1001 to implement the video processing method provided in the method embodiments of this application.
[0147] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.
[0148] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0149] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0150] Display screen 1005 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, disposed on the front panel of terminal 1000; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal 1000. Furthermore, display screen 1005 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0151] The camera assembly 1006 is used to acquire images or videos. In some embodiments, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0152] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.
[0153] The power supply 1008 is used to power the various components in the terminal 1000. The power supply 1008 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0154] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0155] When electronic devices are provided as servers, Figure 11This is a block diagram illustrating a server 1100 according to an exemplary embodiment. The server 1100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The memories 1102 store at least one line of program code, which is loaded and executed by the processor 1101 to implement the video processing methods provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1100 may also include other components for implementing device functions, which will not be elaborated here.
[0156] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1002 or a memory 1102 including instructions. These instructions can be executed by the processor 1001 of the terminal 1000 or the processor 1101 of the server 1100 to complete the aforementioned video processing method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0157] A computer program product includes a computer program / instructions that, when executed by a processor, implement the aforementioned video processing method.
[0158] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0159] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that, The method includes: In response to a video generation instruction for an item, a video generated based on multiple item images and item description text of the item is displayed in the video interface. The multiple item images and the item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to showcase the item from different angles or scenes, and the item description text is used to describe the characteristics of the item. In response to a publishing instruction for the video, the video is published.
2. The video processing method according to claim 1, characterized in that, In response to a video generation instruction for an item, the video interface displays a video generated based on multiple item images and item description text of the item, including at least one of the following: In response to the video generation operation for the item in the video interface, a video generation instruction is generated, and in response to the video generation instruction, a video generated based on multiple item images and item description text of the item is displayed; If the item meets the conditions for generating item content, a video generation instruction is generated, and in response to the video generation instruction, a video generated based on multiple item images and item description text of the item is displayed.
3. The video processing method according to claim 2, characterized in that, The conditions for generating the content of the item include at least one of the following: The inventory quantity of the items is not lower than the quantity threshold; The item is a hot item at the current time, and the hot item is used to indicate the item that is generally needed at the current time. The item has a target identifier, which is used to indicate to an object that possesses the item that needs to distribute the item. The number of searches for keywords related to the item has reached the threshold.
4. The video processing method according to claim 1, characterized in that, The video includes a first split-screen, a second split-screen, and a third split-screen. The first split-screen includes the multiple object images, the second split-screen includes close-up images of the objects among the multiple object images, and the third split-screen includes application images of the objects among the multiple object images. The close-up images are used to demonstrate the features of the objects, and the application images are used to demonstrate the objects in application scenarios. The method further includes: In the video interface, a video of the item is played; During the playback of the video, the first split-screen, the second split-screen, and the third split-screen are displayed sequentially.
5. The video processing method according to claim 4, characterized in that, The video has multiple second-shot frames, each used to display different close-up images, which in turn showcase the features of the item from different angles. The method further includes: During the playback of the video, multiple second-shot scenes are displayed sequentially, and the display size of the close-up images in each second-shot scene is larger than the display size of the images of each item in the first-shot scene.
6. The video processing method according to claim 4, characterized in that, The method further includes: During the playback of the video, the item description text is displayed in the first split-screen, and the display of the item description text is turned off in the second and third split-screens.
7. The video processing method according to claim 1, characterized in that, The method further includes: During the playback of the video, the item description text is displayed as a sticker.
8. The video processing method according to claim 1, characterized in that, The multiple item images include a digital human, which is used to help display the features of the item, and the digital human has the same appearance in each item image in the video.
9. The video processing method according to claim 1, characterized in that, The method further includes: When the playback status of the video meets preset conditions, an item status image is displayed, which is used to show the real-time information of the item.
10. The video processing method according to claim 1, characterized in that, The method further includes: The video content template used to generate the video is displayed. The video content template is used to indicate the narrative structure of the video in which the item is presented, and the video content template matches the category of the item.
11. A video processing method, characterized in that, The method includes: In response to a video generation instruction for an item, image prompts and text prompts are generated based on a single reference image and details of the item. The image prompts are used to indicate the content of the item image in the video, and the text prompts are used to indicate the content of the item description text in the video. Based on the image prompts and text prompts, an artificial intelligence model is used to process the reference image and the details information to obtain multiple item images and item description text for the item. The multiple item images are used to display the item from different angles or scenes. Based on the multiple item images and the item description text, a video for showcasing the items is generated.
12. The video processing method according to claim 11, characterized in that, The generation of image prompts and text prompts based on a single reference image and details of the item includes: Based on the reference image and details of the item, the category of the item is determined; Based on the item's category, the reference image, and the details, the image prompt and the text prompt are generated.
13. The video processing method according to claim 12, characterized in that, The generation of the image prompt and the text prompt based on the item's category, the reference image, and the details includes: Based on the category of the item, a video content template matching the item is determined, and the video content template is used to indicate the narrative structure of the video in displaying the item; Based on the video content template, the reference image, and the details, a video script is generated, which indicates the content of the video to be generated. Based on the video script, the image prompts and the text prompts are generated.
14. The video processing method according to claim 13, characterized in that, The method further includes: If the video content template is a narrative template with a digital human, then character prompts are generated. The digital human is used to assist in displaying the characteristics of the item, and the character prompts are used to indicate that the image of the digital human in each item image in the video remains consistent. The process of processing the reference image based on the image prompts using an artificial intelligence model to obtain multiple image images of the item includes: Based on the image prompts and the character prompts, the reference image is processed using an artificial intelligence model to obtain multiple item images of the item, and the multiple item images include the digital person with a consistent appearance.
15. The video processing method according to claim 11, characterized in that, Before generating the plurality of item images based on the reference image, the method further includes: The reference image is subjected to super-resolution processing, and the main body of the item is extracted from the super-resolution processed reference image to obtain the target image; Based on the image prompts and text prompts, an artificial intelligence model is used to process the reference image and the details information to obtain multiple item images and item description texts for the item, including: Based on the image prompts and text prompts, an artificial intelligence model is used to process the target image and the details information to obtain multiple item images and item description texts for the item.
16. The video processing method according to claim 11, characterized in that, The process of obtaining an item image of the item by processing the reference image based on the image prompts using an artificial intelligence model includes: When the length of the image prompt reaches a length threshold, an item base map is generated based on the image prompt and the artificial intelligence model. The item in the reference image is then transferred to the item base map to obtain an item image of the item. The item base map is used to indicate the scene in which the item is located.
17. A video processing apparatus, characterized in that, The device includes: The display unit is configured to execute a video generation instruction in response to an item, and in the video interface, display a video generated based on multiple item images and item description text of the item. The multiple item images and the item description text are generated by an artificial intelligence model based on a single reference image and detailed information of the item. The multiple item images are used to display the item from different angles or scenes, and the item description text is used to describe the characteristics of the item. The publishing unit is configured to publish the video in response to a publishing instruction for the video.
18. A video processing apparatus, characterized in that, The device includes: The first generation unit is configured to execute a video generation instruction in response to an item, generating image prompts and text prompts based on a single reference image and details of the item, wherein the image prompts are used to indicate the content of the item image in the video, and the text prompts are used to indicate the content of the item description text in the video; The first processing unit is configured to perform processing on the reference image and the details information based on the image prompts and the text prompts using an artificial intelligence model to obtain multiple item images and item description text of the item, wherein the multiple item images are used to display the item from different angles or scenes; The second generation unit is configured to generate a video for displaying the items based on the plurality of item images and the item description text.
19. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the video processing method as described in any one of claims 1 to 16.
20. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the video processing method as described in any one of claims 1 to 16.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video processing method according to any one of claims 1 to 16.