Video generation method and apparatus, electronic device, and storage medium
By receiving a video generation request, determining the type of the target product based on its description information, and obtaining the product video interaction information associated with this type, the target video is generated. This solves the problems of long time and high cost in video generation in the existing technology, and achieves efficient and accurate video generation.
Patent Information
- Application Number
- CN202411605201.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-11
AI Technical Summary
The process of generating spoken-word videos in the existing technology is time-consuming and costly, and has low generation efficiency.
By receiving a video generation request, determining the type of the target product based on its description information, and obtaining product video interaction information associated with the type, the target video is generated.
It shortens the video generation time, reduces costs, and improves generation efficiency and accuracy.
Smart Images

Figure CN119484956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular to the technical field of artificial intelligence such as deep learning and large model, and specifically relates to a video generation method and device, an electronic device and a storage medium. BACKGROUND
[0002] Most of the current oral video is generated by professional personnel manually writing scripts according to promotion needs and selecting appropriate video materials for splicing, but this method has a long video generation process, low video generation efficiency and high cost. SUMMARY
[0003] The present disclosure aims to at least solve one of the technical problems in the related art to some extent.
[0004] The first aspect of the present disclosure provides a video generation method, comprising:
[0005] receiving a video generation request, wherein the generation request includes first description information of a target product;
[0006] determining a target type to which the target product belongs according to the first description information;
[0007] obtaining first interaction information of product videos associated with the target type;
[0008] generating a target video corresponding to the target product according to the first description information and the first interaction information.
[0009] The second aspect of the present disclosure provides a video generation device, comprising:
[0010] a receiving module configured to receive a video generation request, wherein the generation request includes first description information of a target product;
[0011] a determining module configured to determine a target type to which the target product belongs according to the first description information;
[0012] an obtaining module configured to obtain first interaction information of product videos associated with the target type;
[0013] a generating module configured to generate a target video corresponding to the target product according to the first description information and the first interaction information.
[0014] The third aspect of the present disclosure provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video generation method according to the first aspect of the present disclosure is implemented.
[0015] The fourth embodiment of the present disclosure proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the video generation method proposed in the first embodiment of the present disclosure is implemented.
[0016] The fifth embodiment of the present disclosure proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the video generation method proposed in the first embodiment of the present disclosure.
[0017] The video generation method, device, electronic device, and storage medium provided by the present disclosure have the following beneficial effects:
[0018] In the disclosed embodiment, a video generation request is first received. Then, based on the first description information, a target type of a target product is determined. Then, first interaction information of a product video associated with the target type is obtained. Finally, based on the first description information and the first interaction information, a target video corresponding to the target product is generated. Thus, after receiving the video generation request, the target product type is determined based on the product description information. Then, based on the product description information and the interaction information of the product videos associated with the type, a video corresponding to the target product is generated. This improves the efficiency and accuracy of video generation while shortening the video generation time and reducing the cost of video generation.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0021] Figure 1 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0022] Figure 2 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0023] Figure 3 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0024] Figure 4 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0025] Figure 5 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0026] Figure 6 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0027] Figure 7 A flowchart of a video generation method provided by an embodiment of the present disclosure;
[0028] Figure 8 A schematic diagram of the structure of a video generating device provided in an embodiment of the present disclosure;
[0029] Figure 9 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] The present disclosure relates to the fields of artificial intelligence technologies such as deep learning and large models.
[0032] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.
[0033] Deep learning (DL) is the process of learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.
[0034] The large model can also be called the Foundation Model. The model extracts knowledge from billions of corpora or images, learns, and then produces a large model with billions of parameters.
[0035] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0036] The video generation method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0037] Figure 1 A flowchart of a video generation method provided by an embodiment of the present disclosure is shown.
[0038] As shown in Figure 1 , the video generation method can include the following steps:
[0039] Step 101, receiving a video generation request, wherein the generation request includes first description information of a target product.
[0040] It should be noted that the target product can be a product to be generated video, which can be any type of product. For example, the target product can be an electronic device product, a daily necessity product, etc., which is not limited by the present disclosure.
[0041] Among them, the first description information can include any information related to the target product. For example, when the target product is a mobile phone, the first description information can include the target product's business point, i.e. mobile phone, and selling point such as cost performance, etc., which is not limited by the present disclosure.
[0042] In the present disclosure, after receiving the video generation request, the target product to be generated video and the first description information of the target product can be determined, thereby providing conditions for improving the accuracy of generating the video of the target product.
[0043] Step 102, determining the target type to which the target product belongs according to the first description information.
[0044] It should be noted that the target product can be different, and the target type to which it belongs can be different or can be the same. For example, when the target product is a mobile phone, the target type to which it belongs can be a mobile phone type or an electronic device type, and when the target product is a tablet computer, the target type to which it belongs can be a tablet computer type or an electronic device type, which is not limited by the present disclosure.
[0045] In the present disclosure, in order to improve the accuracy and reliability of the target product video generation, after receiving the video generation request, the target type to which the product belongs can be determined based on the first description information of the target product.
[0046] Step 103, obtaining first interaction information of the product video associated with the target type.
[0047] Among them, the first interaction information can be any interaction information. For example, the first interaction information can be to display product information and links to the user when the user watches the video, etc., which is not limited by the present disclosure.
[0048] It should be noted that the target type associated product video can be different, and the corresponding first interaction information can be different, which is not limited by the present disclosure.
[0049] Optionally, in the case of determining the target type to which the product belongs according to the first description information, in order to improve the comprehensiveness and reliability of the obtained first interaction information, the first interaction information of the product video associated with the target type in a certain period of time before the current time can be obtained, and the present disclosure does not make any limitation in this regard.
[0050] Among them, the certain period of time can be pre-set, or can also be determined according to actual conditions, and the present disclosure does not make any limitation in this regard.
[0051] Optionally, in the case of determining the target type to which the product belongs according to the first description information, the first interaction information of each product video in the video information library associated with the target type can also be obtained, and the present disclosure does not make any limitation in this regard.
[0052] Among them, the video information library can be a database for storing product videos associated with the target type, which can be pre-set, and the present disclosure does not make any limitation in this regard.
[0053] In the present disclosure, in the case of determining the target type to which the product belongs according to the first description information, the first interaction information of the product video associated with the target type in a certain period of time before the current time can be obtained according to the need, or the first interaction information of each product video in the video information library associated with the target type can be obtained, thereby improving the accuracy of the determined first interaction information and improving the reliability of the video generation.
[0054] Step 104, generating a target video corresponding to the target product according to the first description information and the first interaction information.
[0055] Among them, the target video can be the final product video corresponding to the target product.
[0056] In the present disclosure, after obtaining the first interaction information of the product video associated with the target type, the target video corresponding to the target product can be generated according to the first description information and the first interaction information of the target product, thereby shortening the time-consuming of generating the video and improving the efficiency and accuracy of the video generation.
[0057] In the embodiments of the present disclosure, first, a video generation request is received, then the target type to which the target product belongs is determined according to the first description information, after that the first interaction information of the product video associated with the target type is obtained, and finally the target video corresponding to the target product is generated according to the first description information and the first interaction information. Thus, after receiving the video generation request, the type to which the target product belongs is determined based on the product description information, and the video corresponding to the target product is generated according to the product description information and the interaction information of the product video associated with the type, thereby improving the efficiency and accuracy of the video generation on the basis of shortening the video generation time and reducing the cost of the video generation.
[0058] Figure 2 A flowchart of a video generation method provided by an embodiment of the present disclosure is shown.
[0059] As shown in Figure 2 , the video generation method can include the following steps:
[0060] Step 201, receiving a video generation request, wherein the generation request includes first description information of a target product.
[0061] Step 202, determining a target type to which the target product belongs according to the first description information.
[0062] Step 203, obtaining first interaction information of product videos associated with the target type.
[0063] The specific implementation forms of steps 201 to 203 can refer to the detailed description in other embodiments of the present disclosure, and will not be repeated here.
[0064] Step 204, determining second description information of each product video according to the first interaction information corresponding to the product video, wherein the second description information includes at least one of the following: interaction frequency, conversion rate.
[0065] The interaction frequency can be the number of interactions of the product video.
[0066] The conversion rate can be the proportion of users who watch the video and complete conversion (such as purchasing the product).
[0067] In the present disclosure, after obtaining the first interaction information of the product videos associated with the target type, the second description information of each product video can be determined according to the first interaction information corresponding to the product video, thereby providing conditions for improving the reliability of video generation.
[0068] Step 205, determining at least one target reference video according to each second description information.
[0069] The target reference video can be a reference video used to generate a target video of the target product.
[0070] In the present disclosure, after determining the second description information of each product video according to the first interaction information corresponding to the product video, at least one target reference video can be determined based on the second description information of each product video. For example, when the second description information includes interaction frequency and conversion rate, the product video with higher interaction frequency and / or conversion rate can be determined as the target reference video, which is not limited by the present disclosure.
[0071] Step 206: Determine reference video elements and reference text content based on at least one target reference video.
[0072] The reference video element may be an element related to the target reference video. For example, the reference video element may be a video frame of the target reference video, or a component in the video frame such as a subject type (e.g., a product, a person, etc.), a color element, etc., which is not limited in this disclosure.
[0073] The reference text content may be the script content of the target reference video, or it may be the spoken copy content, and this disclosure does not limit this.
[0074] In the present disclosure, after determining at least one target reference video, in order to improve the accuracy of the generated video corresponding to the target product, reference video elements and reference text content can be determined based on the at least one target reference video.
[0075] Step 207: Generate a target video based on the first description information, the reference video element, and the reference text content.
[0076] In the present disclosure, after determining the reference video elements and reference text content corresponding to at least one target reference video, a target video can be generated based on the first description information of the target product, the reference video elements and the reference text content, thereby improving the accuracy and reliability of the generated target video.
[0077] In the embodiment of the present disclosure, a video generation request is first received, and then the target type to which the target product belongs is determined based on the first description information, and the first interaction information of the product video associated with the target type is obtained. Then, based on the first interaction information corresponding to each product video, the second description information of the product video is determined, and based on each second description information, at least one target reference video is determined. Finally, based on the at least one target reference video, the reference video element and the reference text content are determined, and the target video is generated based on the first description information, the reference video element and the reference text content. Thus, after receiving the video generation request, based on the description information of the target product, the interaction information of the product video associated with the type to which the target product belongs is obtained, the description information of the product video is determined based on the interaction information, and then the reference video is determined based on the description information. Based on the target product description information, the video elements and the text content of the reference video, the target video of the target product is generated, thereby improving the efficiency and reliability of video generation.
[0078] Figure 3 A flowchart of a video generation method provided by an embodiment of the present disclosure is provided.
[0079] like Figure 3 As shown, the video generation method may include the following steps:
[0080] Step 301, receiving a video generation request, wherein the generation request includes first description information of a target product and video material.
[0081] The video material can be provided by the user, and the specific content thereof can be determined by the user as needed. For example, when the target product is a mobile phone, the video material can be a video material about the performance of the mobile phone, and the like, which is not limited in the present disclosure.
[0082] Step 302, determining a target type to which the target product belongs according to the first description information.
[0083] Step 303, obtaining first interaction information of a product video associated with the target type.
[0084] The specific implementation forms of steps 301 to 303 can refer to the detailed description in other embodiments of the present disclosure, which will not be repeated here.
[0085] Step 304, segmenting the video material to obtain a plurality of material segments.
[0086] It should be noted that when the video material is segmented, the segmentation can be based on the video transition (i.e., video element) in the video material, that is, the segmentation is performed whenever the video subject in the video material changes, which is not limited in the present disclosure.
[0087] It should be noted that the video duration corresponding to each material segment obtained can be different, which is not limited in the present disclosure.
[0088] In the present disclosure, after obtaining the first interaction information of the product video associated with the target type, the video material can be segmented to obtain a plurality of material segments corresponding to the material video, thereby providing a data basis for generating the video of the target product.
[0089] Step 305, determining third description information corresponding to each material segment.
[0090] The third description information can include any information related to the material segment. For example, the third description information can include content description and time length of the material segment, and the like, which is not limited in the present disclosure.
[0091] In the present disclosure, after obtaining the plurality of material segments corresponding to the video material, the third description information corresponding to each material segment can be determined, thereby determining the related information of each material segment and improving the accuracy of video generation.
[0092] Step 306, generating a target video corresponding to the target product according to the first description information, the first interaction information and the third description information.
[0093] In the present disclosure, after obtaining the third description information of each material segment corresponding to the video material, a target video corresponding to the target product can be generated based on the first description information, the first interaction information and the third description information, thereby improving the content accuracy and reliability of the generated target video.
[0094] The specific implementation of step 306 can be referred to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.
[0095] In the disclosed embodiment, a video generation request is first received. Then, based on the first description information, the target type of the target product is determined, and the first interaction information of the product video associated with the target type is obtained. The video material is then segmented to obtain multiple material segments, and the third description information corresponding to each material segment is determined. Finally, based on the first description information, the first interaction information, and the third description information, a target video corresponding to the target product is generated. Thus, after receiving the video generation request, the target product's type is determined based on its description information, and the interaction information of the product video associated with the type is obtained. The video material is segmented, and the description information corresponding to each obtained material segment is determined. Based on the target product description information, the interaction information, and the description information of each material segment, a video of the target product is generated, thereby improving the efficiency of video generation.
[0096] Figure 4 A flowchart of a video generation method provided by an embodiment of the present disclosure is provided.
[0097] like Figure 4 As shown, the video generation method may include the following steps:
[0098] Step 401: Receive a video generation request, wherein the generation request includes first description information of a target product and video material.
[0099] Step 402: Determine the target type of the target product according to the first description information.
[0100] Step 403: Acquire first interaction information of a product video associated with the target type.
[0101] Step 404: Segment the video material to obtain multiple material segments.
[0102] Step 405: Determine the third description information corresponding to each material clip.
[0103] The specific implementation of steps 401 to 405 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be described in detail here.
[0104] Step 406 : Generate a text sequence and a video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information.
[0105] The text sequence may be a script sequence of a video corresponding to a target product, which may include any number of texts, and the present disclosure does not limit this.
[0106] In the present disclosure, after obtaining the third description information corresponding to each material clip, the first description information, the first interaction information and the third description information can be input into a preset large model, and the first description information, the first interaction information and the third description information can be feature fused to obtain the text sequence and the corresponding video sequence corresponding to the target product, thereby improving the efficiency of video generation. The present disclosure does not limit this.
[0107] The preset macro model may be a macro model of any type and structure. For example, the preset macro model may be a multimodal macro model, which is not limited in the present disclosure.
[0108] Step 407 : When the video segment corresponding to any text segment in the text sequence is empty, the duration of the video segment to be generated is determined based on the correspondence between the lengths of other text segments and the durations of the corresponding video segments, and the length of the text segment.
[0109] In the present disclosure, after obtaining the text sequence and the corresponding video sequence corresponding to the target product, when the video segment corresponding to any text segment in the text sequence is empty, in order to accurately and reliably determine the video segment corresponding to any text segment, the correspondence between the length of other text segments and the corresponding video segment duration can be first determined, and based on the correspondence and the length of any text segment, the duration of the video segment to be generated can be determined.
[0110] It should be noted that the longer the text segment is, the longer the corresponding video segment may be, and this disclosure does not limit this.
[0111] Step 408: Generate a video segment corresponding to any text segment based on the length of the text segment and the video segment.
[0112] In the present disclosure, after determining the length of the video segment to be generated, any text segment and the length of the video segment can be input into a preset generation model to obtain the video segment corresponding to the any text segment, thereby improving the accuracy and efficiency of generating the video segment corresponding to the any text segment.
[0113] Step 409: fuse the video segments corresponding to each text segment to obtain a target video.
[0114] In the present disclosure, after generating the video segment of any text segment, the video segments corresponding to each text segment may be fused and rendered to obtain the target video, thereby improving the efficiency of generating the target video.
[0115] In the embodiment of the present disclosure, a video generation request is first received, and then the target type of the target product is determined based on the first description information, and the first interaction information of the product video associated with the target type is obtained. Thereafter, the video material is segmented to obtain multiple material segments, and the third description information corresponding to each material segment is determined. Based on the first description information, the first interaction information and the third description information, a text sequence and a corresponding video sequence corresponding to the target product are generated. When the video segment corresponding to any text segment in the text sequence is empty, the duration of the video segment to be generated is determined based on the correspondence between the length of other text segments and the duration of the corresponding video segment, and the length of any text segment. Finally, based on any text segment and the duration of the video segment, a video segment corresponding to any text segment is generated, and the video segments corresponding to each text segment are fused to obtain the target video. Therefore, based on the target product description information, the interactive information of the product video associated with the type of the target product and the description information of each video material segment, the text sequence and the corresponding video sequence corresponding to the target product are obtained. When the video segment corresponding to any text segment is empty, the length of the video segment corresponding to any text segment is determined according to the correspondence between the length of other text segments and the corresponding video segment market, and the corresponding video segment is generated. The video segments corresponding to each text segment are fused to obtain the target video corresponding to the target product, thereby realizing the automatic generation of product videos, reducing the cost of video generation, and improving the efficiency of video generation.
[0116] Figure 5 A flowchart of a video generation method provided by an embodiment of the present disclosure is provided.
[0117] like Figure 5 As shown, the video generation method may include the following steps:
[0118] Step 501: Receive a video generation request, wherein the generation request includes first description information of a target product and video material.
[0119] Step 502: Determine the target type of the target product according to the first description information.
[0120] Step 503: Acquire first interaction information of a product video associated with the target type.
[0121] Step 504: Segment the video material to obtain multiple material segments.
[0122] Step 505, determine the third description information corresponding to each material segment.
[0123] Step 506, generate the text sequence and the corresponding video sequence of the target product according to the first description information, the first interaction information and the third description information.
[0124] The specific implementation forms of steps 501 to 506 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be repeated here.
[0125] Step 507, in the case that the video segment corresponding to any text segment in the text sequence is empty, obtain the reference text segment adjacent to the any text segment in the text sequence and the reference video segment corresponding to the reference text segment, respectively.
[0126] In the present disclosure, after obtaining the text sequence and the corresponding video sequence of the target product, in the case that the video segment corresponding to any text segment in the text sequence is empty, in order to make the video segment of the determined any text segment content consistent with the video segments corresponding to other text segments, the reference text segment adjacent to the any text segment in the text sequence and the reference video segment corresponding to the reference text segment can be obtained first.
[0127] Step 508, generate the video segment corresponding to any text segment according to the any text segment, the reference text segment and the reference video segment.
[0128] In the present disclosure, after determining the reference text segment and the reference video segment adjacent to the any text segment, the any text segment, the reference text segment and the reference video segment can be input into the generation model, so that the model can generate the video segment corresponding to the any text segment based on the context text segment information and the context video segment information of the any text segment, thereby improving the accuracy and reliability of the generated video segment.
[0129] Step 509, fuse the video segment corresponding to each text segment to obtain the target video.
[0130] In the present disclosure, after generating the video segment corresponding to the any text segment based on the any text segment, the reference text segment and the reference video segment, the video segment corresponding to each text segment is fused to obtain the target video, thereby ensuring the content coherence and consistency of the target video and improving the reliability and efficiency of the generated target video.
[0131] The specific implementation form of step 509 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be repeated here.
[0132] In the embodiments of the present disclosure, first, a video generation request is received, then a target type to which a target product belongs is determined according to first description information, first interaction information of a product video associated with the target type is obtained, after that, video materials are segmented to obtain a plurality of material segments, third description information corresponding to each material segment is determined, and a text sequence corresponding to the target product and a video sequence corresponding to the target product are generated according to the first description information, the first interaction information and the third description information, in the case that a video segment corresponding to any text segment in the text sequence is empty, a reference text segment adjacent to the any text segment in the text sequence and reference video segments corresponding to the reference text segment are obtained, and a video segment corresponding to the any text segment is generated according to the any text segment, the reference text segment and the reference video segments, and finally, the video segments corresponding to each text segment are fused to obtain a target video. Therefore, after obtaining the text sequence corresponding to the target product and the video sequence corresponding to the target product based on the target product description information, the product video interaction information associated with the target type to which the target product belongs and the description information of each video material segment, in the case that the video segment corresponding to any text segment is empty, the adjacent text segment and the corresponding video segment of the any text segment are obtained, and the video segment corresponding to the any text segment is generated based on the any text segment, the adjacent text segment and the corresponding video segment, the video segments corresponding to each text segment are fused to obtain the target video corresponding to the target product, thereby improving the efficiency and accuracy of video generation.
[0133] Figure 6 A flowchart of a video generation method provided by an embodiment of the present disclosure.
[0134] As shown in Figure 6 , the video generation method can include the following steps:
[0135] Step 601, receiving a video generation request, wherein the first description information of the target product is included in the generation request.
[0136] Step 602, determining the target type to which the target product belongs according to the first description information.
[0137] Step 603, obtaining the first interaction information of the product video associated with the target type.
[0138] Step 604, generating the target video corresponding to the target product according to the first description information and the first interaction information.
[0139] The specific implementation forms of steps 601 to 604 can refer to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.
[0140] Step 605, playing the target video, wherein the playing interface of the target video includes a modification control.
[0141] In the present disclosure, after generating the target video corresponding to the target product, the target video can be played so that the user can confirm the effect and content of the generated target video.
[0142] It should be noted that the modification control can be any type of control. For example, the modification control can be a button type control, and the present disclosure does not limit this.
[0143] Step 606, in the case where it is monitored that the modification control is triggered, a modification page is displayed, wherein the modification page includes at least one of the following: a first control for triggering text content modification, and a second control for triggering modification of a video frame.
[0144] It should be noted that the type of the first control can be the same as the type of the second control, or can also be different from the type of the second control, and the present disclosure does not limit this.
[0145] In the present disclosure, after playing the target video, in the case where it is monitored that the modification control in the playing interface of the target video is triggered, a modification page can be displayed to the user, so that the user can modify the text content and / or video frame of the target video, and make the target video more in line with the user's needs.
[0146] Step 607, according to the editing instruction received in the modification page, the target video is updated.
[0147] In the present disclosure, after displaying the modification page, the target video can be updated according to the editing instruction received in the modification page, so that the target video is more in line with the user's needs, and the personalization of the target video is improved.
[0148] It should be noted that the editing instruction can include at least one of an editing instruction for the text content of the target video and an editing instruction for the video frame, and can also include specific editing content of the user, and the present disclosure does not limit this.
[0149] In the embodiments of the present disclosure, first, a video generation request is received, then a target type to which a target product belongs is determined according to first description information, and first interaction information of a product video associated with the target type is obtained, then a target video corresponding to the target product is generated according to the first description information and the first interaction information, finally the target video is played, and in the case that a modification control is triggered, a modification page is displayed, and the target video is updated according to an editing instruction received in the modification page. Therefore, after the target video of the target product is generated according to the description information of the target product and the interaction information of the product video associated with the type to which the target product belongs, the target video is played, and in the case that the modification control in the playing interface of the target video is triggered, the modification page is displayed, the editing instruction in the modification page is received, and the target video is updated, thereby improving the personalization and accuracy of the generated video.
[0150] Figure 7 A flowchart of a video generation method provided by an embodiment of the present disclosure.
[0151] As shown in Figure 7 , the video generation method can include the following steps:
[0152] Step 701, receiving a video generation request, wherein the first description information of the target product is included in the generation request.
[0153] Step 702, determining the target type to which the target product belongs according to the first description information.
[0154] Step 703, obtaining first interaction information of a product video associated with the target type.
[0155] Step 704, generating a target video corresponding to the target product according to the first description information and the first interaction information.
[0156] The specific implementation forms of steps 701 to 704 can refer to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.
[0157] Step 705, in the case that a target video publishing instruction is received, publishing the target video.
[0158] In the present disclosure, after the target video corresponding to the target product is generated, the target video can be published in the case that a target video publishing instruction is received, thereby realizing the promotion of the target product.
[0159] Step 706, obtaining second interaction information of the target video.
[0160] The second interaction information can include at least one of the interaction frequency and the conversion rate of the target video, which is not limited in the present disclosure.
[0161] In the present disclosure, after the target video is published, in order to determine the promotion effect of the target video on the target product, the second interaction information of the target video can be acquired to determine the interaction feedback of the target video.
[0162] In step 707, the target video is updated when the second interaction information meets the first condition.
[0163] The first condition can be an interaction information condition for determining whether to update the target video, which can be pre-set or determined according to actual conditions, and can include any condition. For example, the first condition can be that the target video is updated when the interaction frequency is less than or equal to the interaction frequency threshold value, and / or the conversion rate is less than or equal to the conversion rate threshold value, and the present disclosure does not limit this.
[0164] In the present disclosure, after obtaining the second interaction information of the target video, when the second interaction information meets the first condition, it can be determined that the interaction information feedback of the target video is poor, and the promotion effect of the target video on the target product is not good. At this time, the target video needs to be updated to improve the accuracy and reliability of the target video.
[0165] Optionally, when the first condition includes the frequency threshold value and the conversion rate threshold value, when the interaction frequency in the second interaction information is less than or equal to the frequency threshold value, and / or the conversion rate is less than or equal to the conversion rate threshold value, it can be determined that the promotion effect of the target video on the target product during the publishing period is poor. At this time, the third interaction information of the product video associated with the target type during the publishing period of the target video can be acquired first, and then the target video is updated according to the third interaction information, so as to improve the quality of the target video and improve the promotion effect on the target product.
[0166] The frequency threshold value can be an interaction frequency threshold value for determining whether to update the target video, which can be pre-set or determined according to actual conditions, and the present disclosure does not limit this.
[0167] The conversion rate threshold value can be a conversion rate threshold value for determining whether to update the target video, which can be pre-set or determined according to actual conditions, and the present disclosure does not limit this.
[0168] It should be noted that the third interaction information can be different from the first interaction information, and the present disclosure does not limit this.
[0169] Optionally, when the first condition includes a frequency threshold and a conversion rate threshold, when the interaction frequency in the second interaction information is greater than the frequency threshold, and / or the conversion rate is greater than the conversion rate threshold, it can be determined that the interaction information feedback of the target video is relatively good, and the target video has a good promotion effect on the target product during the release period. At this time, the target video and the second interaction information can be associated and stored in a target type-associated video information library, so that the target type-associated video information library stores product videos with good interaction information feedback, ensuring the real-time and reliability of the data in the target type-associated video information library.
[0170] In the disclosed embodiment, a video generation request is first received. Then, based on the first description information, the target type of the target product is determined, and the first interaction information of the product video associated with the target type is obtained. Then, based on the first description information and the first interaction information, a target video corresponding to the target product is generated. Upon receiving a target video release instruction, the target video is released. Finally, second interaction information of the target video is obtained, and the target video is updated if the second interaction information satisfies the first condition. Thus, after generating the target video corresponding to the target product based on the target product description information and the product video interaction information associated with its type, upon receiving a release instruction, the target video is released, and if the target video's interaction information satisfies the condition, the target video is updated, thereby improving the quality of the generated target video and improving the reliability of video generation.
[0171] In order to implement the above embodiments, the present disclosure also proposes a video generating device.
[0172] Figure 8 This is a structural diagram of a video generation device provided in an embodiment of the present disclosure.
[0173] like Figure 8 As shown, the video generating device 800 includes: a receiving module 801 , a determining module 802 , an acquiring module 803 , and a generating module 804 .
[0174] The receiving module 801 is configured to receive a video generation request, wherein the generation request includes first description information of a target product;
[0175] A determination module 802 is configured to determine a target type to which the target product belongs based on the first description information;
[0176] An acquisition module 803 is configured to acquire first interaction information of a product video associated with a target type;
[0177] The generating module 804 is configured to generate a target video corresponding to the target product according to the first description information and the first interaction information.
[0178] In a possible implementation of the present disclosure, the generating module 804 is specifically configured to:
[0179] Determine second description information of the product video based on the first interaction information corresponding to each product video, wherein the second description information includes at least one of the following: interaction frequency and conversion rate;
[0180] Determining at least one target reference video according to each second description information;
[0181] Determining reference video elements and reference text content based on at least one target reference video;
[0182] A target video is generated according to the first description information, the reference video element and the reference text content.
[0183] In a possible implementation of the present disclosure, the generation request further includes video material, and the generation module 804 is further configured to:
[0184] Split the video material into multiple clips;
[0185] Determining third description information corresponding to each material clip;
[0186] A target video corresponding to the target product is generated according to the first description information, the first interaction information, and the third description information.
[0187] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0188] Generate a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information;
[0189] When the video segment corresponding to any text segment in the text sequence is empty, the duration of the video segment to be generated is determined based on the correspondence between the lengths of other text segments and the durations of the corresponding video segments, and the length of any text segment;
[0190] Generate a video clip corresponding to any text clip based on the length of any text clip and video clip;
[0191] The video segments corresponding to each text segment are fused to obtain the target video.
[0192] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0193] Generate a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information;
[0194] When the video segment corresponding to any text segment in the text sequence is empty, obtain reference text segments adjacent to any text segment in the text sequence, and reference video segments corresponding to the reference text segments respectively;
[0195] Generate a video segment corresponding to any text segment based on any text segment, reference text segment and reference video segment;
[0196] The video segments corresponding to each text segment are fused to obtain the target video.
[0197] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0198] Playing the target video, wherein the playback interface of the target video includes modification controls;
[0199] In the case where it is detected that the modification control is triggered, a modification page is displayed, wherein the modification page includes at least one of the following: a first control for triggering modification of text content, and a second control for triggering modification of a video frame;
[0200] The target video is updated according to the editing instructions received in the modification page.
[0201] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0202] Upon receiving a target video publishing instruction, publishing the target video;
[0203] Obtaining second interaction information of the target video;
[0204] When the second interaction information satisfies the first condition, the target video is updated.
[0205] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0206] When the interaction frequency in the interaction information is less than or equal to the frequency threshold, and / or the conversion rate is less than or equal to the conversion rate threshold, obtaining third interaction information of product videos associated with the target type during the release period of the target video;
[0207] The target video is updated according to the third interaction information.
[0208] In a possible implementation of the present disclosure, the generating module 804 is further configured to:
[0209] When the interaction frequency in the second interaction information is greater than the frequency threshold, and / or the conversion rate is greater than the conversion rate threshold, the target video and the second interaction information are associated and stored in a target type associated video information library.
[0210] In a possible implementation of the present disclosure, the generation module 804 is further configured to perform any one of the following:
[0211] obtain the first interaction information of the product video associated with the target type within a certain time period before the current time;
[0212] obtain the first interaction information of each product video in the video information library associated with the target type.
[0213] The functions and specific implementation principles of the above modules in the embodiments of the present disclosure can be referred to the above method embodiments, which will not be described here.
[0214] In the embodiments of the present disclosure, first, a video generation request is received, then the target type to which the target product belongs is determined according to the first description information, after that, the first interaction information of the product video associated with the target type is obtained, and finally, the target video corresponding to the target product is generated according to the first description information and the first interaction information. Therefore, after receiving the video generation request, the type to which the target product belongs is determined based on the product description information, and the video corresponding to the target product is generated according to the product description information and the interaction information of the product video associated with the type, thereby improving the efficiency and accuracy of video generation on the basis of shortening the video generation time and reducing the cost of video generation.
[0215] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0216] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0217] As Figure 9As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0218] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0219] The computing unit 901 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the video generation method by any other appropriate means (e.g., by means of firmware).
[0220] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0221] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable video generation apparatus to produce a machine, such that the program code, when executed by the processor or controller, causes the machine to perform functions / operations specified in the flowcharts and / or block diagrams. The program code can execute entirely on a machine, partly on a machine, as a stand-alone software package, partly on a machine and partly on a remote machine or entirely on a remote machine or server.
[0222] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0223] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0224] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0225] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0226] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, as long as the desired results of the technical solutions of the present disclosure can be achieved.
[0227] Furthermore, the terms "first", "second", etc. are used only for descriptive purposes and do not connote or imply relative importance or a quantity of the indicated technical features. Thus, a feature defined with "first", "second", etc. can include at least one of the features implicitly or explicitly. In the description of the disclosure, the meaning of "a plurality" is at least two, for example, two, three, etc., unless otherwise specifically defined. In the description of the disclosure, the words "if" and "when" can be interpreted as "at the time of" or "when" or "in response to determining" or "in the case of".
[0228] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A video generation method, comprising: receiving a video generation request, wherein the generation request includes first description information of a target product; Determining the target type of the target product according to the first description information; Obtaining first interaction information of a product video associated with the target type, wherein the first interaction information is displaying product information and a link to the user when the user watches the video; Generating a target video corresponding to the target product according to the first description information and the first interaction information; wherein generating a target video corresponding to the target product according to the first description information and the first interaction information includes: Determine second description information of the product video based on the first interaction information corresponding to each product video, wherein the second description information includes at least one of the following: interaction frequency and conversion rate; When the second description information includes interaction frequency and conversion rate, a product video with a higher interaction frequency and / or conversion rate is determined as a target reference video; Determining reference video elements and reference text content based on the target reference video; The target video is generated according to the first description information, the reference video element and the reference text content.
2. The method according to claim 1, wherein The generation request further includes video material, and generating a target video corresponding to the target product according to the first description information and the first interaction information includes: Segmenting the video material to obtain multiple material segments; Determining third description information corresponding to each of the material clips; A target video corresponding to the target product is generated according to the first description information, the first interaction information, and the third description information.
3. The method according to claim 2, wherein: Generating a target video corresponding to the target product according to the first description information, the first interaction information, and the third description information includes: generating a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information; When the video segment corresponding to any text segment in the text sequence is empty, determining the duration of the video segment to be generated based on the correspondence between the lengths of other text segments and the durations of the corresponding video segments, and the length of any text segment; Generate a video segment corresponding to the text segment according to the text segment and the duration of the video segment; The video segments corresponding to each text segment are fused to obtain the target video.
4. The method according to claim 2, wherein: Generating a target video corresponding to the target product according to the first description information, the first interaction information, and the third description information includes: generating a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information; When the video segment corresponding to any text segment in the text sequence is empty, obtaining reference text segments adjacent to the text segment in the text sequence and reference video segments corresponding to the reference text segments respectively; generating a video segment corresponding to the any text segment according to the any text segment, the reference text segment, and the reference video segment; The video segments corresponding to each text segment are fused to obtain the target video.
5. The method according to claim 1, wherein After generating the target video corresponding to the target product, the method further includes: Playing the target video, wherein the playback interface of the target video includes a modification control; In the case where it is detected that the modification control is triggered, a modification page is displayed, wherein the modification page includes at least one of the following: a first control for triggering text content modification, and a second control for triggering modification of a video frame; The target video is updated according to the editing instruction received in the modification page.
6. The method of claim 1, wherein: After generating the target video corresponding to the target product, the method further includes: Upon receiving a target video publishing instruction, publishing the target video; Acquire second interaction information of the target video; When the second interaction information satisfies the first condition, the target video is updated.
7. The method according to claim 6, wherein: When the interaction information satisfies the first condition, updating the target video includes: When the interaction frequency in the second interaction information is less than or equal to a frequency threshold, and / or the conversion rate is less than or equal to a conversion rate threshold, obtaining third interaction information of product videos associated with the target type during the release of the target video; The target video is updated according to the third interaction information.
8. The method of claim 6, wherein: After obtaining the second interaction information of the target video, the method further includes: When the interaction frequency in the second interaction information is greater than a frequency threshold, and / or the conversion rate is greater than a conversion rate threshold, the target video and the second interaction information are associated and stored in a target type-associated video information library.
9. The method of claim 8, wherein: The acquiring of the first interaction information of the product video associated with the target type includes any one of the following: Acquire first interaction information of product videos associated with the target type within a certain period before the current moment; The first interaction information of each product video in the video information library associated with the target type is obtained.
10. A video generating device, wherein: The device comprises: A receiving module, configured to receive a video generation request, wherein the generation request includes first description information of a target product; a determination module, configured to determine the target type to which the target product belongs based on the first description information; an acquisition module, configured to acquire first interaction information of a product video associated with the target type, wherein the first interaction information is product information and a link displayed to the user when the user watches the video; a generating module, configured to generate a target video corresponding to the target product according to the first description information and the first interaction information; The generation module is specifically used to: Determine second description information of the product video based on the first interaction information corresponding to each product video, wherein the second description information includes at least one of the following: interaction frequency and conversion rate; When the second description information includes interaction frequency and conversion rate, a product video with a higher interaction frequency and / or conversion rate is determined as a target reference video; Determining reference video elements and reference text content based on the target reference video; The target video is generated according to the first description information, the reference video element and the reference text content.
11. The device according to claim 10, wherein The generation request also includes video material, and the generation module is further configured to: Segmenting the video material to obtain multiple material segments; Determining third description information corresponding to each of the material clips; A target video corresponding to the target product is generated according to the first description information, the first interaction information, and the third description information.
12. The device according to claim 11, wherein The generating module is further configured to: generating a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information; When the video segment corresponding to any text segment in the text sequence is empty, determining the duration of the video segment to be generated based on the correspondence between the lengths of other text segments and the durations of the corresponding video segments, and the length of any text segment; Generate a video segment corresponding to the text segment according to the text segment and the duration of the video segment; The video segments corresponding to each text segment are fused to obtain the target video.
13. The device according to claim 11, wherein The generating module is further configured to: generating a text sequence and a corresponding video sequence corresponding to the target product according to the first description information, the first interaction information, and the third description information; When the video segment corresponding to any text segment in the text sequence is empty, obtaining reference text segments adjacent to the text segment in the text sequence and reference video segments corresponding to the reference text segments respectively; generating a video segment corresponding to the any text segment according to the any text segment, the reference text segment, and the reference video segment; The video segments corresponding to each text segment are fused to obtain the target video.
14. The device according to claim 10, wherein The generating module is further configured to: Playing the target video, wherein the playback interface of the target video includes a modification control; In the case where it is detected that the modification control is triggered, a modification page is displayed, wherein the modification page includes at least one of the following: a first control for triggering text content modification, and a second control for triggering modification of a video frame; The target video is updated according to the editing instruction received in the modification page.
15. The apparatus of claim 10, wherein: The generating module is further configured to: Upon receiving a target video publishing instruction, publishing the target video; Acquire second interaction information of the target video; When the second interaction information satisfies the first condition, the target video is updated.
16. The apparatus of claim 15, wherein: The generating module is further configured to: When the interaction frequency in the interaction information is less than or equal to the frequency threshold, and / or the conversion rate is less than or equal to the conversion rate threshold, obtaining third interaction information of the product video associated with the target type during the release of the target video; The target video is updated according to the third interaction information.
17. The apparatus of claim 15, wherein: The generating module is further configured to: When the interaction frequency in the second interaction information is greater than a frequency threshold, and / or the conversion rate is greater than a conversion rate threshold, the target video and the second interaction information are associated and stored in the target type associated video information library.
18. The apparatus of claim 17, wherein: The generation module is further used for any of the following: Acquire first interaction information of product videos associated with the target type within a certain period before the current moment; The first interaction information of each product video in the video information library associated with the target type is obtained.
19. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that may be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Video generation method, information display method and computing device
CN115908694A
Video generation method and device, equipment and storage medium
CN116389849A