A method, apparatus, device, medium and program product for processing a video list
By first displaying the initial video without special effects in the video list, and then overlaying external special effects onto the initial video, the problem of list display being blocked due to long video generation time is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-10-10
- Publication Date
- 2026-07-21
AI Technical Summary
Videos take much longer to generate than images and text. This causes a slow generation of a single video in the video list to block the display of the entire list, increasing user waiting time and resulting in a poor user experience.
By first displaying placeholder icons in the video list to generate an initial video without special effects, and then displaying it in the video list, external special effects are added and overlaid onto the initial video to achieve progressive enhancement.
It shortens the time from video generation to visibility, avoids the problem of slow generation of a single video blocking the display of the video list, and improves the user experience.
Smart Images

Figure CN119135988B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer technology, and more particularly to a method, apparatus, device, medium, and program product for processing video lists. Background Technology
[0002] Video generation takes significantly longer than image and text generation. If multiple videos are generated for specific materials, the video list can only be displayed for users to view the video effects after all videos in the list have been generated. This method can easily cause a single, slow-generating video to block the display of the entire video list, increasing the time required for other videos to become visible. During this process, users cannot perform other operations, resulting in a poor user experience. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, medium, and program product for processing video lists, which can shorten the time required for the process from video generation to video visibility and improve user experience.
[0004] In a first aspect, embodiments of this disclosure provide a method for processing a video list, including:
[0005] Obtain the target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object;
[0006] The system displays placeholders corresponding to at least one target video in a video list, generates at least one initial video of the target object based on the target content, and replaces the corresponding placeholders in the video list with the successfully generated initial video.
[0007] Based on the target content, special effects materials for at least one initial video are generated, and the special effects materials are superimposed onto the initial videos in the video list according to preset layout information to obtain the target video.
[0008] Secondly, embodiments of this disclosure also provide a video list processing apparatus, the apparatus comprising:
[0009] The content acquisition module is used to acquire target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object;
[0010] The list display module is used to display placeholders corresponding to at least one target video through a video list, generate at least one initial video of the target object based on the target content, and replace the corresponding placeholders in the video list with the successfully generated initial video;
[0011] The video overlay module is used to generate special effects material for at least one initial video based on the target content, and to overlay the special effects material onto the initial video in the video list according to preset layout information to obtain the target video.
[0012] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0013] One or more processors;
[0014] Storage device for storing one or more programs.
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the video list processing method as described in any embodiment of this disclosure.
[0016] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video list processing method as described in any embodiment of this disclosure.
[0017] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video list processing method as described in any embodiment of this disclosure.
[0018] This disclosure provides a method, apparatus, device, medium, and program product for processing a video list. By acquiring the target content of a target object, at least one target video corresponding to the target object is generated. A placeholder identifier corresponding to at least one target video is displayed in the video list. First, at least one initial video of the target object is generated based on the target content. The successfully generated initial video replaces the corresponding placeholder identifier in the video list. This avoids the problem of a single, slowly generated video blocking the display of the entire video list, shortening the time from video generation to video visibility. During the generation and display of the initial video, special effects materials for generating at least one initial video based on the target content are executed in parallel. The successfully generated special effects materials are superimposed onto the initial video according to preset layout information to obtain the target video. This disclosure achieves progressive enhancement by first generating an initial video without special effects materials and displaying it in the video list, and then superimposing the successfully generated special effects materials onto the initial video using an external plugin. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 A flowchart illustrating a video list processing method provided in an embodiment of this disclosure;
[0021] Figure 2 A schematic diagram of a video list provided in an embodiment of this disclosure;
[0022] Figure 3 This is a schematic diagram of a preset layout information provided in an embodiment of the present disclosure;
[0023] Figure 4 A flowchart illustrating another video list processing method provided in this embodiment of the disclosure;
[0024] Figure 5 This is a schematic diagram of a video rendering process provided in an embodiment of the present disclosure;
[0025] Figure 6 A schematic diagram of a video list processing device provided in an embodiment of this disclosure;
[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0038] Figure 1This is a flowchart illustrating a video list processing method provided in an embodiment of the present disclosure. This embodiment is applicable to scenarios involving the generation of video lists, such as generating a video list of product demonstration videos. The method can be executed by a video list processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.
[0039] like Figure 1 As shown, the method includes:
[0040] S110. Obtain the target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object.
[0041] The target object can represent any object in the video. The video can be a short video or other types of video. For example, in an e-commerce scenario, the target object can be a product, service, or tourist attraction. Target content refers to content associated with the target object. Specifically, target content includes target object information and target object attributes. Target object information represents the target object from the perspective of object content. For example, target object information can include object images and descriptive text. Target object attributes represent the target object from the perspective of object attributes. Target object attributes can include object type or object characteristics.
[0042] In this embodiment of the disclosure, the target content may include product information, service information, or tourist attraction information, etc. Product information can be obtained through product links or product detail pages, etc. Similarly, service information can be obtained through service links or service detail pages, etc. The target video can represent a video generated for a target object. For example, the target video can be an explanatory video of the target object. Typically, target videos have a marketing nature, capable of showcasing the characteristics of the target object from multiple aspects, dimensions, and angles to attract users to make purchases.
[0043] For example, obtaining the target content corresponding to the target object includes: obtaining the target page corresponding to the target object based on the target URL; and parsing the target page to obtain the target content. The target page can represent a webpage carrying the target content of the target object. In some embodiments, the target URL is obtained through a user interaction page. If a generation operation is detected on the user interaction page, the corresponding target page is obtained based on the target URL. The target page is then parsed to obtain the target content. For example, the URL of product A on e-commerce platform A, entered in the user interaction page, is obtained. The corresponding target page is obtained based on the URL of product A, and the target page is parsed to obtain the name, selling points information, and images of product A. Selling points information is usually a characteristic of product A, or attribute information of other products.
[0044] S120. Display placeholders corresponding to the at least one target video in the video list, generate at least one initial video of the target object based on the target content, and replace the placeholders in the video list with the successfully generated initial video.
[0045] The video list is a display list of generated videos. It displays at least one target video generated for the target object. Optionally, placeholders can be displayed in the video list to indicate the display position of each target video. These placeholders can be layout diagrams of the target videos. For example, they can be rectangles, etc., and this embodiment does not specifically limit this. The initial video can be a video synthesized based on target content through steps such as material matching, cover cropping, and title extraction. Material matching can refer to selecting image content from the target content as material for video synthesis. Cover cropping can refer to the operation of generating a video cover by combining the material and preset cover layout information. Title extraction can refer to the operation of generating a video title based on the selling points or promotional information of the target object in the target content.
[0046] If a user enters a target URL on the user interaction page and clicks the generate button, a video list will be displayed on the user interaction page. The video list includes at least one placeholder for each target video, indicating the display position of each target video. Figure 2 This is a schematic diagram illustrating a video list provided in an embodiment of this disclosure. For example... Figure 2 As shown, placeholders 210 corresponding to four target videos are displayed in the video list 200, and a loading message is displayed at the position corresponding to the placeholders 210.
[0047] It should be noted that in the initial stage, since no initial video has been successfully generated, only a loading prompt message may be displayed on the user interaction page. If any initial video is successfully generated, placeholders for the initial video and other target videos will be displayed in the video list.
[0048] Alternatively, if no initial video is successfully generated, a corresponding number of placeholder markers can be displayed in the video list based on the number of target videos. This disclosure does not specifically limit this approach.
[0049] For example, target material is determined based on the target object information; a video cover is generated based on the target material and a preset cover layout; a title script is generated based on the target object information; a video title is generated based on the title script; at least one initial video of the target object is generated based on the target material, the video cover, and the video title. The placeholder identifier in the video list is replaced with the successfully generated initial video.
[0050] In this embodiment, the target object information can represent the content of the target object. The target material can represent the material used to generate the video. The material includes images and text, etc. In an e-commerce scenario, the target material can include the main product image, product detail images, product selling point images, and product scene images, etc. The video cover can be generated according to a preset cover layout combined with the matched material. Optionally, video frames that accurately reflect the video content and video title can also be extracted from the video cover set as video covers. The video cover set includes candidate video frames extracted from the initial video at fixed time points to serve as video covers.
[0051] Video titles reflect the content of the video, helping users quickly understand it. Title scripts can represent different focuses of the target audience, generated based on information about the target audience. These focuses may include promotions or selling points. Since different focus titles represent the target audience from different dimensions, a title script with a particular focus can be determined based on the central idea or core of the initial video. The video title is then generated based on the content of the title script.
[0052] Because target video generation involves steps such as material matching, subtitle generation, virtual object overlay, cover cropping, title extraction, and video compositing, the process is time-consuming. Furthermore, the video list is only displayed after all target videos have been generated, and since it contains several target videos, the display can easily be blocked by a single, slower-processing target video. For example, with a video list of ten target videos, the above video generation process takes approximately one and a half minutes. Users must wait for this process to complete, during which time they cannot perform other operations, resulting in a poor user experience. Analysis shows that the time-consuming operations in the above video generation process include subtitle generation and virtual object overlay. Therefore, these time-consuming operations can be extracted from the overall video generation process. After material matching, cover cropping, and title extraction, video compositing can be performed to obtain the initial video. Then, the initial video is rendered into the video list to replace the corresponding placeholders. Figure 2 As shown, after the initial video A corresponding to the placeholder marker 210 located in the upper left corner of the video list 200 is successfully generated, the placeholder marker 210 is replaced with the initial video A to display the video cover 220 of the initial video A in the video list. Other display markers 210 display the message "Generation in progress". At this time, the initial video A may not include special effects materials. If the initial video B corresponding to another placeholder marker 210 is successfully generated, the newly generated initial video B is used to replace the corresponding placeholder marker 210.
[0053] Optionally, the initial video is played in response to an interactive action on the initial video displayed in the video list. As long as the initial video is displayed in the video list, the user can click on it to view the generated video effect.
[0054] S130. Generate special effects material for at least one initial video based on the target content, and overlay the special effects material onto the initial video in the video list according to preset layout information to obtain the target video.
[0055] The special effects materials can be the processing results of steps in the target video generation process that meet preset time conditions. In this embodiment, the special effects materials include subtitles generated by the subtitle generation step and / or virtual objects generated by the virtual object matching step. Virtual objects can represent objects in the video used to explain the target object. For example, virtual objects can include digital humans or virtual cartoon characters. Preset layout information characterizes the layout of the special effects materials in the video. Specifically, the preset layout information includes subtitle position, subtitle style, virtual object position information, and virtual object timestamp information. For example, the preset layout information can include subtitle position, subtitle color, font, background color, virtual object position, and virtual object timestamp information. Figure 3 This is a schematic diagram illustrating a preset layout information provided in an embodiment of this disclosure. Figure 3 As shown, the layout information of each video frame can be defined based on the timeline, including: 0-5 seconds is the first type of layout style 310, 6-10 seconds is the second type of layout style 320, 10-20 seconds is the third type of layout style 330, etc.
[0056] Since the initial videos have already been generated and displayed in the video list, once the subtitles and / or virtual objects for any of the initial videos are generated, they can be overlaid onto the corresponding initial videos as add-ons to achieve progressive enhancement. At this point, the target video can include subtitles and / or virtual objects. For example, the target video could be a product demonstration video that includes subtitles and a digital human.
[0057] For example, the virtual object's image and audio track are determined based on the target object's attributes; a virtual object for the initial video is generated based on the image and audio track; a subtitle script is generated based on the target object information; and subtitle content and timestamp information are generated based on the subtitle script and the audio track, wherein the timestamp information represents the timestamp of the initial video frame corresponding to the subtitle content. The virtual object is used to explain the target object, and the explanation content is consistent with the subtitle content. For example, in an e-commerce scenario, a digital human explains the product. Or, in a game explanation scenario, a digital human explains the game, etc.
[0058] The subtitle position, subtitle style, virtual object position information, and virtual object timestamp information in the initial video are determined according to the preset layout information; the virtual object is superimposed onto the initial video in the video list according to the virtual object position information and virtual object timestamp information; subtitles are generated according to the subtitle content and subtitle style, and the subtitles are superimposed onto the initial video in the video list according to the subtitle position and subtitle timestamp information.
[0059] In some embodiments, an object image is matched from candidate digital humans based on the attributes of the product. For example, if the product is product x, a digital human with the image c that matches product x can be matched. Then, an object audio track is matched from candidate audio tracks based on the digital human with image c. Since the audio track reflects the audio playback speed, if the subtitle content is determined, the playback duration required to play the subtitle content can be determined based on the object audio track. Optionally, an object image can also be matched from candidate digital humans based on the product type. For example, if product x is a beauty product, a digital human with one image is matched from the candidate digital humans. If product x is kitchenware, a digital human with another image is matched from the candidate digital humans.
[0060] In some embodiments, the subtitle script can represent different types of subtitle content. The subtitle script can be generated based on target object information and a preset template. The preset template can be based on commonly used product description content. The corresponding subtitle content for the video type is determined from the different types of subtitle scripts based on the video title. For example, promotional videos require promotional subtitles, or videos introducing product selling points require product selling point description subtitles, etc. After determining the subtitle content, the number of characters in the subtitle content can be obtained. The subtitle playback speed of the virtual object (e.g., the number of characters played per unit time) is determined based on the object's audio track. The playback duration of the subtitle content is determined by combining the number of characters in the subtitle content and the subtitle playback speed. Then, the subtitle timestamp information is determined based on the video duration of the initial video corresponding to the special effects material and the playback duration of the subtitle content.
[0061] In some embodiments, the subtitle position information includes the subtitle's position relative to the video, including top and left distances, used for subtitle positioning. The subtitle style includes subtitle metadata, such as subtitle color, background color, or font, used for rendering the subtitle. Optionally, the corresponding subtitle position and style can be matched from candidate subtitle layouts based on target object information and target object attributes. Alternatively, different types of target objects can be matched with different subtitle positions and styles based on different video titles. This disclosure does not specifically limit this. The virtual object layout includes virtual object position information and virtual object timestamp information, such as virtual object position information and virtual object timestamp information. For example, virtual object position information and virtual object timestamp information can be matched from candidate virtual object layouts based on different video titles.
[0062] The technical solution of this disclosure involves obtaining the target content of a target object to generate at least one target video corresponding to the target object. A placeholder identifier corresponding to at least one target video is displayed in a video list. First, at least one initial video of the target object is generated based on the target content. The successfully generated initial video replaces the corresponding placeholder identifier in the video list. This avoids the problem of a single, slowly generated video blocking the display of the entire video list, shortening the time from video generation to video visibility. During the generation and display of the initial video, special effects materials for generating at least one initial video based on the target content are executed in parallel. The successfully generated special effects materials are overlaid onto the initial video according to preset layout information to obtain the target video. This disclosure also achieves progressive enhancement by first generating an initial video without special effects materials and displaying it in the video list, and then overlaying the successfully generated special effects materials onto the initial video using an external plugin.
[0063] Figure 4 This is a flowchart illustrating another video list processing method provided by an embodiment of this disclosure. Based on the above embodiments, this disclosure further defines the processing flow for special effects materials. For example... Figure 4 As shown, the method includes:
[0064] S410. Obtain the target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object, and the target content includes target object attributes and target object information.
[0065] S420. Display placeholders corresponding to the at least one target video in the video list, generate at least one initial video of the target object based on the target content, and replace the placeholders in the video list with the successfully generated initial video.
[0066] S430. Determine the object image and object audio track of the virtual object based on the target object attributes, and generate the virtual object of the initial video based on the object image and object audio track.
[0067] For example, candidate virtual avatars are matched based on the target object attributes to obtain the object image of the virtual object in the initial video. Candidate object audio tracks are matched based on the object image to obtain the object audio track of the virtual object.
[0068] Multiple candidate virtual images are predefined. Then, the object image of the virtual object is matched from the candidate virtual images based on the attributes of the target object. Specifically, the object image of the virtual object is matched from the candidate virtual images based on the product type. For example, if the product is an eraser, the product type can be determined as stationery, and the image of a student can be matched from the candidate images as the object image.
[0069] By matching candidate object audio tracks with the object image, the virtual object's object audio track is obtained. Based on the virtual object's object audio track, information such as timbre, volume, and playback speed can be determined.
[0070] S440. Generate a subtitle script based on the target field in the target object information, and generate subtitle content based on the subtitle script, wherein the target field represents information in the target object information that affects the conversion amount of the target object.
[0071] In this embodiment, the subtitle script may include different types of product description text, including product description text presented in the first person or product description text presented in the third person. A subtitle script explaining the selling points can be generated by combining formulaic phrases and the product selling point field in the target audience information. Then, subtitle content is extracted from the subtitle script explaining the selling points. Optionally, a subtitle script explaining promotions can also be generated by combining formulaic promotional phrases and the product promotion field in the target audience information. Then, subtitle content is extracted from the subtitle script explaining promotions.
[0072] S450. Determine the playback duration of the subtitle content based on the object audio track, and determine the subtitle timestamp information based on the playback duration and the initial video duration.
[0073] The number of words that can be played per unit time can be determined based on the audio track and subtitle content. Then, based on the subtitle content and the number of words played per unit time, the playback duration of the subtitle content can be determined. Since the subtitles should also finish playing after the video playback is complete, the timestamp of the initial video frame corresponding to the subtitle content can be determined based on the playback duration of the subtitle content and the initial video duration.
[0074] S460. Determine the subtitle position, subtitle style, virtual object position information, and virtual object timestamp information in the initial video based on the preset layout information.
[0075] S470. Based on the virtual object's location information and timestamp information, the virtual object is superimposed onto the initial video in the video list.
[0076] S480. Generate subtitles based on the subtitle content and style, and overlay the subtitles onto the initial video in the video list based on the subtitle position and subtitle timestamp information.
[0077] Since the display positions of each target video in the video list are indicated by placeholders, if the initial video at a certain position is successfully generated, the placeholder at that position is replaced by the successfully generated initial video. Then, if the special effects material corresponding to any initial video in the video list has been rendered, the special effects material is overlaid on the initial video as an external plugin to complete the progressive enhancement.
[0078] Figure 5 This is a schematic diagram of a video rendering process provided in an embodiment of the present disclosure. Since the subtitle generation and virtual object overlay processes are relatively time-consuming during video rendering, such as... Figure 5 As shown, material matching (510), title extraction (520), and cover cropping (530) can be performed first. After material matching, title extraction, and cover cropping are completed, video compositing (540) can be performed to obtain an initial video and render it to the video list. During this process, the subtitle generation and virtual object matching steps continue to be performed to determine the virtual object position information (550), virtual object meta-information (560), and virtual object timestamp information (570), and render the virtual object based on the virtual object meta-information information (560); and to determine the subtitle position information (580), subtitle meta-information information (590), and subtitle timestamp information (5100), and render the subtitle content based on the subtitle meta-information information (590). If the virtual object corresponding to the initial video has been rendered, the rendered virtual object is superimposed on the corresponding video frame of the initial video in the video list according to the virtual object position information (550) and virtual object timestamp information (570). If the subtitle corresponding to the initial video has been rendered, the rendered subtitle is superimposed on the corresponding video frame of the initial video in the video list according to the subtitle position and subtitle timestamp information.
[0079] The subtitle metadata includes subtitle content, color, background color, and font, used to render the subtitle content. Subtitle position information refers to the subtitle's position relative to the video frame, including top and left distances, used to locate the subtitle. Subtitle timestamp information refers to the video timestamp corresponding to the current subtitle content, used to determine when to render this subtitle. Virtual object metadata includes the object's appearance and audio track, used to render the virtual object. Virtual object position information refers to the virtual object's position relative to the video frame, including top and left distances, used to locate the virtual object. Virtual object timestamp information refers to the video timestamp corresponding to the current virtual object, used to determine when to render this virtual object.
[0080] S490. Display the target video in the video list.
[0081] S4100. If the initial video generation fails, a retry control is displayed in the video list at the placeholder corresponding to the failed initial video. The retry control is used to indicate that the video generation failed and to trigger a regeneration operation.
[0082] Reference Figure 2 In the initial stage, all videos are in a compositing state, and the user interface displays a loading prompt. If any initial video is successfully composited, it is displayed in the video list 200, while other unsuccessfully composited videos are displayed as placeholders 210. The compositing status of unsuccessfully composited videos is monitored in real time. When a new initial video is successfully composited, it replaces the placeholder 210 in the corresponding position in the video list 200. If video compositing fails, a retry control 230 is displayed at the corresponding placeholder 210.
[0083] The technical solution of this disclosure extracts the less time-consuming material matching, title extraction, and cover cropping processes during video rendering, synthesizes an initial video after these processes are completed, and renders the initial video to a video list for display, thus shortening the time required for video to become visible. During the above process, the more time-consuming subtitle rendering and virtual object rendering processes are executed in parallel. After the subtitles and virtual objects are rendered, they are overlaid onto the corresponding initial video in the video list, achieving progressive enhancement of the videos in the video list during the video generation process.
[0084] Figure 6 This is a schematic diagram of a video list processing device provided in an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0085] like Figure 6 As shown, the device includes: a content acquisition module 610, a list display module 620, and a video overlay module 630.
[0086] The content acquisition module 610 is used to acquire target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object;
[0087] The list display module 620 is used to display placeholders corresponding to at least one target video through a video list, generate at least one initial video of the target object according to the target content, and replace the corresponding placeholders in the video list with the successfully generated initial video;
[0088] The video overlay module 630 is used to generate special effects material for at least one initial video based on the target content, and overlay the special effects material onto the initial video in the video list according to preset layout information to obtain the target video.
[0089] Optionally, the content acquisition module 610 is specifically used for:
[0090] Obtain the target page corresponding to the target object based on the target URL;
[0091] The target page is parsed to obtain the target content.
[0092] Optionally, the target content includes target object information;
[0093] The list display module 620 is specifically used for:
[0094] The target material is determined based on the target object information, and a video cover is generated based on the target material and a preset cover layout.
[0095] A title script is generated based on the target object information, and a video title is generated based on the title script.
[0096] Generate at least one initial video of the target object based on the target material, video cover, and video title.
[0097] Optionally, the target content includes target object attributes and target object information; the special effects materials include virtual objects and subtitles;
[0098] The video overlay module 630 is specifically used for:
[0099] The virtual object's image and audio track are determined based on the target object's attributes, and the virtual object of the initial video is generated based on the object's image and audio track.
[0100] A subtitle script is generated based on the target object information. Subtitle content and subtitle timestamp information are generated based on the subtitle script and the object audio track. The subtitle timestamp information represents the timestamp of the initial video frame corresponding to the subtitle content.
[0101] Further, the step of generating a subtitle script based on the target object information, and generating subtitle content and subtitle timestamp information based on the subtitle script and the object audio track, includes:
[0102] A subtitle script is generated based on the target field in the target object information, and subtitle content is generated based on the subtitle script. The target field represents the information in the target object information that affects the conversion rate of the target object.
[0103] The playback duration of the subtitle content is determined based on the audio track of the object, and the subtitle timestamp information is determined based on the playback duration and the initial video duration.
[0104] Furthermore, the video overlay module 630 is specifically used for:
[0105] The subtitle position, subtitle style, virtual object position information, and virtual object timestamp information in the initial video are determined based on the preset layout information;
[0106] Based on the location information and timestamp information of the virtual object, the virtual object is superimposed onto the initial video in the video list;
[0107] Subtitles are generated based on the subtitle content and style, and then superimposed onto the initial video in the video list based on the subtitle position and timestamp information.
[0108] Optionally, it also includes:
[0109] The failure retry module is used to display a retry control in the video list at the placeholder corresponding to the failed initial video if the initial video generation fails. The retry control is used to indicate that the video generation failed and to trigger a regeneration operation.
[0110] Optionally, it also includes:
[0111] A video playback module is used to play the initial video in response to an interactive operation on the initial video displayed in the video list.
[0112] The video list processing apparatus provided in this disclosure can execute the video list processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0113] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0114] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0115] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.
[0116] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0117] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0118] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0119] The electronic device provided in this embodiment and the video list processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0120] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video list processing method provided in the above embodiments.
[0121] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0124] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0125] Obtain the target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object;
[0126] The system displays placeholders corresponding to at least one target video in a video list, generates at least one initial video of the target object based on the target content, and replaces the corresponding placeholders in the video list with the successfully generated initial video.
[0127] Based on the target content, special effects materials for at least one initial video are generated, and the special effects materials are superimposed onto the initial videos in the video list according to preset layout information to obtain the target video.
[0128] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0130] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0131] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0132] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0134] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0135] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for processing a video list, characterized in that, include: Obtain target content corresponding to the target object, wherein the target content is used to generate at least one target video corresponding to the target object, and the target content includes target object attributes and target object information; the target object information represents the target object from the object content dimension, including object images and object description text, and the target object attributes represent the target object from the object attribute dimension, including object type or object characteristics; The system displays placeholders corresponding to at least one target video in a video list, generates at least one initial video of the target object based on the target content, and replaces the corresponding placeholders in the video list with the successfully generated initial video. In response to an interactive operation on the initial video displayed in the video list, the initial video is played; During the generation and display of the initial video, the virtual object's image and audio track are determined based on the target object's attributes. The virtual object of the initial video is generated based on the object image and audio track. A subtitle script is generated based on the target object information. Subtitle content and subtitle timestamp information are generated based on the subtitle script and the audio track. The virtual object and subtitle content are superimposed onto the initial video in the video list according to preset layout information to obtain the target video. The subtitle timestamp information represents the timestamp of the initial video frame corresponding to the subtitle content.
2. The method according to claim 1, characterized in that, The step of obtaining the target content corresponding to the target object includes: Obtain the target page corresponding to the target object based on the target URL; The target page is parsed to obtain the target content.
3. The method according to claim 1, characterized in that, The target content includes target object information; The step of generating at least one initial video of the target object based on the target content, and replacing the corresponding placeholder in the video list with the successfully generated initial video, includes: The target material is determined based on the target object information, and a video cover is generated based on the target material and a preset cover layout. A title script is generated based on the target object information, and a video title is generated based on the title script. Generate at least one initial video of the target object based on the target material, video cover, and video title.
4. The method according to claim 1, characterized in that, The step of generating a subtitle script based on the target object information, and generating subtitle content and subtitle timestamp information based on the subtitle script and the object audio track, includes: A subtitle script is generated based on the target field in the target object information, and subtitle content is generated based on the subtitle script. The target field represents the information in the target object information that affects the conversion rate of the target object. The playback duration of the subtitle content is determined based on the audio track of the object, and the subtitle timestamp information is determined based on the playback duration and the initial video duration.
5. The method according to claim 1, characterized in that, The step of overlaying the virtual object and subtitle content onto the initial video in the video list according to preset layout information includes: The subtitle position, subtitle style, virtual object position information, and virtual object timestamp information in the initial video are determined based on the preset layout information; Based on the location information and timestamp information of the virtual object, the virtual object is superimposed onto the initial video in the video list; Subtitles are generated based on the subtitle content and style, and then superimposed onto the initial video in the video list based on the subtitle position and timestamp information.
6. The method according to claim 1, characterized in that, Also includes: If the initial video generation fails, a retry control is displayed in the video list at the placeholder corresponding to the failed initial video. The retry control is used to indicate that the video generation failed and to trigger a regeneration operation.
7. A video list processing apparatus, characterized in that, include: The content acquisition module is used to acquire target content corresponding to the target object. The target content is used to generate at least one target video corresponding to the target object. The target content includes target object attributes and target object information. The target object information represents the target object from the object content dimension, including object images and object description text. The target object attributes represent the target object from the object attribute dimension, including object type or object characteristics. The list display module is used to display placeholders corresponding to at least one target video through a video list, generate at least one initial video of the target object based on the target content, and replace the corresponding placeholders in the video list with the successfully generated initial video; A video playback module is used to play the initial video in response to an interactive operation on the initial video displayed in the video list; The video overlay module is used to determine the virtual object's image and audio track based on the target object's attributes during the generation and display of the initial video; generate the virtual object of the initial video based on the object's image and audio track; generate a subtitle script based on the target object information; generate subtitle content and subtitle timestamp information based on the subtitle script and the object's audio track; and overlay the virtual object and subtitle content onto the initial video in the video list according to preset layout information to obtain the target video. The subtitle timestamp information represents the timestamp of the initial video frame corresponding to the subtitle content.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video list processing method as described in any one of claims 1-6.
9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the processing method for the video list as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video list processing method as described in any one of claims 1-6.