A video generation method, apparatus, electronic device, storage medium, and program
By acquiring the correlation information between historical short videos and long videos, generating location information, and extracting long video footage, the problem of incomplete plots or slow pacing in historical short videos is solved, thus achieving high-quality short video generation and distribution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING IQIYI TECH CO LTD
- Filing Date
- 2024-08-21
- Publication Date
- 2026-08-04
AI Technical Summary
Within video platforms, historical short videos often have incomplete storylines or slow pacing, resulting in poor appeal to unfamiliar users and making effective promotion difficult.
By acquiring the correlation information between the target's historical short videos and preset long videos, and using the correlation information between the images and dialogue to generate location information, the long video material range is extracted to generate new videos to complete the plot.
It generates short video materials with complete storylines and fast pace, which are suitable for distribution to unfamiliar users on fast-paced platforms, thus improving the promotional effect.
Smart Images

Figure CN119052602B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia processing technology, and in particular to a video generation method, a video generation apparatus, an electronic device, a computer-readable storage medium, and a computer program. Background Technology
[0002] Long videos generally refer to videos that are more than half an hour long, mainly films and TV dramas, and are usually produced by professional companies; short videos, also known as short films, are a form of internet content dissemination, generally videos with a duration of less than 5 minutes that are disseminated on new media platforms.
[0003] Within video platforms, long videos can typically be edited into numerous short videos during historical promotional periods. These short videos can correspond to the same or different key plot points of the long videos, serving as promotional material for the long videos. However, if the initial purpose of producing these short videos is for in-platform advertising or fan interaction, and given the long production time and the lack of access to the original long video's media information, incomplete plots or slow pacing in the short videos can make them less appealing to unfamiliar users. This lack of understanding of the narrative context can hinder effective promotion and distribution to new audiences. Summary of the Invention
[0004] The purpose of this invention is to provide a video generation method, a video generation device, an electronic device, a computer-readable storage medium, and a computer program to supplement the narrative information of historical short videos on a website, thereby enabling the processing of historical short videos on the website. The specific technical solutions are as follows:
[0005] In a first aspect of this invention, a video generation method is provided, the method comprising:
[0006] Obtain target historical short videos, and obtain the association information between the target historical short videos and the long video materials corresponding to preset long videos; the association information includes image association information and dialogue association information;
[0007] Based on the image association information and the dialogue association information, the location information of the target historical short video in the preset long video is generated; the location information is used to indicate the video range of at least one long video material.
[0008] The preset long video is cropped based on the location information, and a new video is generated based on the cropped video segment.
[0009] In a second aspect of the invention, a video generation apparatus is also provided, the apparatus comprising:
[0010] The association information acquisition module is used to acquire target historical short videos and acquire association information between the target historical short videos and the long video materials corresponding to preset long videos; the association information includes scene association information and dialogue association information.
[0011] The location information acquisition module is used to generate location information of the target historical short video in the preset long video based on the image association information and the dialogue association information; the location information is used to indicate the video range of at least one long video material.
[0012] The video generation module is used to extract segments from the preset long video based on the location information, and generate new videos based on the extracted video segments.
[0013] In another aspect of the present invention, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement any of the video generation methods described above when executing the programs stored in the memory.
[0014] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform any of the video generation methods described above.
[0015] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the video generation methods described above.
[0016] The video generation method, apparatus, electronic device, and storage medium provided in this invention obtain the association information between the target historical short video and the corresponding long video material of the preset long video. Based on the aforementioned image association information and dialogue association information, the position information of the target historical short video in the preset long video is determined. The position information can be used to indicate the video range of at least one long video material. At this time, the preset long video can be cropped based on the generated position information to determine the video range of the original long video corresponding to the historical short video on the site, thereby obtaining the complete video plot. A new video is generated based on the cropped video segment to complete the plot information of the historical short video on the site, realizing the processing of the historical short video on the site for promotion and distribution to users. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0018] Figure 1 This is a flowchart illustrating the steps of a video generation method embodiment of the present invention.
[0019] Figure 2 This is a flowchart illustrating the steps of another video generation method embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram illustrating the video generation process provided in an embodiment of the present invention;
[0021] Figure 4 This is a structural block diagram of an embodiment of a video generation device according to the present invention;
[0022] Figure 5 This is a schematic diagram of the structure of an electronic device embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0024] Within video platforms, during historical promotional periods, long videos can typically be edited to produce a large number of short videos. Different short videos can correspond to the same or different highlights of the long video, which can then be used to promote the long video.
[0025] This invention addresses the issue of incomplete or slow-paced short videos produced on the platform's historical video content. It utilizes video analysis and understanding technology to process selected video segments or short videos from the past, automatically generating short video materials suitable for distribution to users on fast-paced short video platforms. Specifically, to address the problem of lost original media or basic information of historical videos, it combines image content retrieval, dialogue analysis, and dialogue matching technologies to automatically generate long video segment information corresponding to the historical short videos. This involves determining the location information of the target historical short video within a preset long video based on image and dialogue association information, thereby extracting segments from the preset long video to provide basic materials for the production of new short videos. Furthermore, video analysis technology is used to process and synthesize selected segments of the long video, generating a fast-paced, multi-cut short video with a complete and engaging storyline, suitable for promotion and distribution to unfamiliar users.
[0026] Reference Figure 1 The diagram illustrates a flowchart of a video generation method embodiment of the present invention, which may specifically include the following steps:
[0027] Step 101: Obtain the target historical short video and obtain the association information between the target historical short video and the long video material corresponding to the preset long video;
[0028] The target historical short video can refer to short videos or selected video clips produced in the past on the video site. This invention can process such videos. The preset long video can refer to existing videos on the video site, such as data from several movies or several episodes of TV series. Since such videos are usually produced by professional companies and uploaded to the video site, they mainly serve as the original medium when editing short videos. They contain complete plot information. In this invention, such videos can provide long video material when processing the target historical short video.
[0029] To facilitate the provision of long video materials in subsequent processing, the association information between the target historical short video and the corresponding long video materials of the preset long video can be obtained at this time. This association information can be used to indicate the degree of association between the target historical short video and each preset long video in the preset video site.
[0030] Step 102: Based on the image association information and dialogue association information, generate the location information of the target historical short video in the preset long video;
[0031] In one embodiment of the present invention, the long video material range of the preset long video corresponding to the target historical short video can be located based on the degree of correlation between the target historical short video and the preset long video. Specifically, the video segments corresponding to the long video material range with a high degree of correlation can be used as video materials that are likely to complete the plot information of the target historical short video.
[0032] Each video segment can contain a pair of start / end point information. At this time, the point information of the target historical short video in the preset long video can be generated to indicate the video interval of the corresponding video segment based on the start time of the video where the start point information is located and the end time of the video where the end point information is located.
[0033] In practical applications, the degree of correlation between the target historical short video and the preset long video can be determined by combining the correlation between visual elements and the correlation between dialogue elements. Therefore, when locating a long video clip range, this can be represented by locating the long video clip range corresponding to the visual elements of the target historical short video, and locating the long video clip range corresponding to the dialogue elements of the target historical short video. In specific implementations, when generating location information based on the correlation information, the location information of the long video corresponding to the target historical short video can be generated by performing consistency checks and fusion of the visual correlation information and the dialogue correlation information. It should be noted that this embodiment of the invention does not limit the method for fusion and verification of correlation information, nor the specific method for generating location information.
[0034] Step 103: Extract a preset long video segment based on the location information, and generate a new video based on the extracted video segment.
[0035] The obtained location information is used to indicate the video segment corresponding to the target historical short video within the video range of the preset long video. The preset long video can provide video material that can complete the plot information of the target historical short video. At this time, the location information can be used to process the target historical short video. The processing can be represented by extracting the video range of the preset long video corresponding to the target historical short video, thereby generating a new video based on the extracted video segment, and automatically generating short video material suitable for distribution to unfamiliar users on fast-paced short video platforms.
[0036] In this embodiment of the invention, by obtaining the association information between the target historical short video and the corresponding long video material of the preset long video, the position information of the target historical short video in the preset long video is determined based on the aforementioned image association information and dialogue association information. The position information can be used to indicate the video range of at least one long video material. At this time, the preset long video can be cropped based on the generated position information to determine the video range of the original long video corresponding to the historical short video on the site, thereby obtaining the complete video plot. A new video is generated based on the cropped video segment to complete the plot information of the historical short video on the site, thereby realizing the processing of the historical short video on the site for promotion and distribution to users.
[0037] Reference Figure 2 The diagram illustrates a flowchart of another video generation method embodiment of the present invention, which may specifically include the following steps:
[0038] Step 201: In the preset video site, locate the long video material range of the preset long video corresponding to the target historical short video;
[0039] In this embodiment of the invention, target historical short videos can be processed to supplement the plot information of historical videos on the site that have incomplete storylines or slow pace, generating mixed short videos with complete and exciting storylines and faster pace. This process enables the processing of historical short videos on the site for promotion and distribution to users.
[0040] In one embodiment of the present invention, the association information between the target historical short video and the preset long video can first be obtained. This association information can be used to indicate the degree of association between the target historical short video and each preset long video in the preset video site, so as to locate the long video material range of the preset long video corresponding to the target historical short video based on the degree of association between the target historical short video and the preset long video.
[0041] In practical applications, the degree of correlation between the target historical short video and the preset long video can be determined by combining the correlation between visuals and dialogue. Therefore, when locating a long video clip range, this can be represented by locating the long video clip range corresponding to the visuals of the target historical short video, and locating the long video clip range corresponding to the dialogue of the target historical short video. That is, the obtained correlation information between the target historical short video and the preset long video can include visual correlation information and dialogue correlation information. Visual correlation information can be used to indicate the first long video clip range corresponding to the visuals of the target historical short video, and dialogue correlation information can be used to indicate the second long video clip range corresponding to the dialogue of the target historical short video.
[0042] The acquisition of image association information can be represented by locating the first long video material range corresponding to the images contained in the target historical short video from a preset video site, and obtaining the image association information.
[0043] Specifically, the degree of correlation between images can be determined by the similarity of the image features in the video.
[0044] In practical implementation, the image features of the target historical short video can be obtained and used as query features for similarity retrieval. Then, a reference image feature library for a preset long video can be obtained. This library can be constructed from the image features of several preset long videos, meaning the image features of the preset long videos can be used as reference features for similarity retrieval. Next, a similarity search can be performed between the image features of the target historical short video and the reference image feature library. If the similarity exceeds a preset first threshold, the target long video material corresponding to the reference image feature is associated with the image corresponding to that image feature. At this point, the video time of the target long video material can be obtained. Based on the video time of the target long video material, a first long video material interval is determined within the preset long videos, obtaining the image association information between the first long video material interval and the image corresponding to the image feature.
[0045] In practical applications, similarity retrieval is performed as image content retrieval, and the extracted image features can be image features. For example... Figure 3 As shown, the image association process can include video frame extraction, image feature extraction, and image content retrieval.
[0046] During the process of video frame extraction and image feature extraction, image feature extraction is performed on the target historical short video. Specifically, key frames are extracted from the target historical short video, and feature expressions of the whole image or key areas of the image are extracted from the corresponding key frames as query features for image content retrieval. Image feature extraction is also performed on the preset long video. Specifically, key frames are extracted from the preset long video, and feature expressions of the whole image or key areas of the image are extracted from the corresponding key frames as reference features for image content retrieval.
[0047] Among them, a keyframe refers to a video frame that can express the content of the video. The keyframe can be extracted by means of equal-interval frame extraction or by means of lens detection and lens keyframe extraction. This embodiment of the invention does not limit the method. The overall image area can include all elements in the image, such as people, objects, background, text, etc. The key area of the image mainly refers to the area corresponding to the key objects in the image, such as the area where the person is located.
[0048] Feature representation can primarily involve processing image regions using existing deep learning models to obtain feature vectors. This can be achieved by using a CLIP model (Contrastive Language-Image Pretraining, a cross-modal large model) to extract features from the overall video frame or key video regions. Alternatively, CLIP can be used to extract overall video frame features, followed by the use of a human recognition model to identify features in the human (key) region. It should be noted that the deep learning model used can be a CLIP-type model or a human recognition model used to extract features from human body regions (key regions); this embodiment of the invention does not impose any limitations on this.
[0049] In the process of image content retrieval, for example, the target historical short video can be a short video produced in the past or a selected video segment within the video site, and the preset long video can be an existing long video on the video site. At this time, each feature extracted from the historical short video, let's say feat_q_i, can be used for similarity retrieval in the reference screen feature library constructed through the features extracted from the preset long video, let's say {feat_r_j}. When the similarity between the two is greater than a preset first threshold th1 (e.g., th1 = 0.8), it can be considered that the short video screen feature feat_q_i and the long video feature feat_r_j, that is, the screen content at time i of the short video and the screen content at time j of the long video are associated.
[0050] Furthermore, the acquisition of dialogue association information can be achieved by locating the second long video material range corresponding to the dialogue contained in the target historical short video from the preset video site, thereby obtaining the dialogue association information.
[0051] Specifically, the degree of relevance between lines can be determined by the similarity of the line information.
[0052] In practice, dialogue analysis can be performed on the target historical short video and the preset long video separately to obtain the first dialogue information of each line in the target historical short video and the second dialogue information of each line in the preset long video. Then, the first dialogue information and the second dialogue information can be similarly calculated. If the similarity between the first dialogue information and the second dialogue information exceeds a preset second threshold, it is determined that the time interval in which the first dialogue information exists is associated with the dialogue content in the time interval in which the second dialogue information exists. At this time, the video time of the associated dialogue content can be obtained. Based on the video time of the associated dialogue content, the second long video material interval is determined in the preset long video to obtain the dialogue association information between the second long video material interval and the associated dialogue content.
[0053] In practical applications, the dialogue information obtained from dialogue analysis can include the text content of the dialogue and the time interval of the dialogue. The time interval of the dialogue can be determined based on the start and end times of the dialogue.
[0054] For example, such as Figure 3 As shown, the dialogue association process can include dialogue analysis and dialogue matching. Assuming the target historical short video can be a short video produced historically on the video platform or a selected video clip, and the preset long video can be an existing long video on the video platform, the dialogue analysis process for the target historical short video can utilize dialogue analysis technology to process the historical short video and extract the dialogue information for each line of dialogue corresponding to the short video, assuming it is {(sub_q_m,sta_tm_q_m,end_tm_q_m)}, which can include the dialogue text content (sub_q_m) and the start and end information of the dialogue (sta_tm_q_m,end_tm_q_m). Similarly, the dialogue analysis process for the preset long video can utilize dialogue analysis technology to process the long video and extract the dialogue information for each line of dialogue corresponding to the long video, assuming it is {(sub_r_n,sta_tm_r_n,end_tm_r_n)}, which can also include the dialogue text content and the start and end information of the dialogue.
[0055] In the process of dialogue matching, dialogue matching technology can be used to calculate the similarity of the dialogue text content of each pair of short-long dialogues in the obtained first and second dialogue information.
[0056] Dialogue matching technology is a method used to analyze and compare the similarity between different lines in a script. First, the BERT text model can be used to extract dialogue text features. Specifically, this involves segmenting the dialogue text into words or phrases, followed by stop word removal, stemming, and other processing. The processed vocabulary is then converted into numerical form. At this point, the cosine similarity between pairwise dialogue vectors of short and long lines can be calculated, or the ratio of the intersection to the union of the short and long line sets can be calculated, or the edit distance can be calculated to obtain the similarity between the first and second line information. The calculated edit distance refers to the minimum number of editing operations required to convert one line into another. Editing during the conversion process includes insertion, deletion, and replacement, which are not limited in this embodiment of the invention.
[0057] In practical applications, when the similarity between the two is greater than the preset second threshold th2 (e.g., th2 = 0.8), it can be considered that the time interval [sta_tm_q_m, end_tm_q_m] in which the short video dialogue m exists is related to the dialogue content in the time interval [sta_tm_r_n, end_tm_r_n] in which the long video dialogue n exists.
[0058] In a preferred embodiment, if the first dialogue information can be matched to obtain multiple second dialogue information, that is, the number of second dialogue information whose similarity is greater than a preset second threshold can be multiple sentences, then multiple matching results can be filtered, and one sentence of second dialogue information can be selected from multiple matching results as the matching result.
[0059] The filtering strategy can be expressed as selecting the sentence with the highest similarity as the associated interval, or using a dynamic programming scheme to map the association results of short and long lines, selecting the matching results that conform to the ascending order of the start time as the associated interval. The conforming to the ascending order of the start time means that, assuming the time pairs of the extracted dialogue information [sta_tm_q_m, end_tm_q_m] from the short video are [3,4][5,6][9,10], the corresponding matching time pairs [sta_tm_r_n, end_tm_r_n] are: [3,4] corresponds to [50,51], [100,102]; [5,6] corresponds to [55,56]; [9,10] corresponds to [59,60]; the start times 3, 5, and 9 conform to the ascending order, so [50,51], [55,56], and [59,60] that conform to the ascending order of 3, 5, and 9 are selected as the output matching results.
[0060] Step 202: Merge the first long video clip interval and the second long video clip interval to obtain the union time interval;
[0061] In one embodiment of the present invention, after locating the long video material range of the preset long video corresponding to the target historical short video based on the degree of correlation between the target historical short video and the preset long video, the point information of the long video corresponding to the target historical short video can be generated by performing consistency verification and fusion of the image association information and the dialogue association information.
[0062] In practical applications, video clips corresponding to long video clip segments with high correlation are used as video clips that are likely to complete the plot information of the target's historical short video. The fusion and verification of the correlation information of the images and the correlation information of the dialogue make up for the problem of incomplete matching of a single modality. This can be mainly manifested as merging the first long video clip segment and the second long video clip segment.
[0063] In the specific implementation, the long video material range represents the video segments of the target historical short video that are highly related to the preset long video. The video range of this video segment is determined based on the video start time and video end time. At this time, when merging the long video material range, it is manifested as merging its time range, and then the union time range of the two is used as the output to obtain the union time range.
[0064] Step 203: Use the union of time intervals as the location information of the target historical short video in the preset long video;
[0065] The obtained union time interval can be used to indicate the final association result of the joint image association information and the dialogue association information. At this time, the complete association point information of the target historical short video in the preset long video can be generated. Since the point information is used to indicate the video interval of the corresponding video segment, the obtained union time interval can be used as the video interval of the corresponding video segment.
[0066] In a preferred embodiment of the present invention, the obtained union time interval can also be filtered to remove isolated segments with too short duration (e.g., less than th3, assuming th3 = 1 second) and generate the fused associated interval.
[0067] Step 204: Extract segments from the preset long video based on the location information, and generate a new video based on the extracted video segments.
[0068] The corresponding video segment is the video segment of the target historical short video within the video segment of the preset long video. At this time, a video segment located in the video segment can be extracted from the preset long video. The extracted video segment can provide the target historical short video with information that can complete the plot. Based on the extracted video segment, a new video is generated, automatically generating short video materials suitable for distribution to unfamiliar users on fast-paced short video platforms.
[0069] Specifically, the video interval can be determined based on the video start time and video end time. At this time, the preset long video can be trimmed based on the video start time and video end time to obtain multiple trimmed video segments. Then, the multiple video segments are spliced together to obtain the generated video.
[0070] In a preferred embodiment of the present invention, a preset long video can be cropped based on the generated fused correlation interval, and multiple cropped video segments can be merged to generate a mixed-cut short video.
[0071] In a preferred embodiment of the present invention, regardless of whether the preset long video is trimmed based on the union time interval or the fused associated interval, the trimming process can also incorporate the combination of dialogue integrity and video content completeness to trim and merge the long video content, generating a mixed-cut short video suitable for content promotion and distribution on existing short video platforms. Specifically, the trimmed video segments must simultaneously meet the requirements of containing complete dialogue duration without truncating individual lines of dialogue, and containing complete shot duration without truncating individual shots. This ensures that the trimmed video segments contain complete narrative semantics, thereby enabling the generation of a mixed-cut short video containing plot information from the trimmed video segments while maintaining the semantic integrity of the newly produced video based on the trimmed segments containing complete narrative semantics.
[0072] In this embodiment of the invention, for short videos produced in the past that have incomplete plots or slow pacing, video analysis and understanding technology can be used to process selected video segments or short videos from the past, automatically generating short video materials suitable for distribution to users on fast-paced short video platforms. Addressing the issue of lost original media or basic information of historical videos, image content retrieval, dialogue analysis, and dialogue matching technologies can be combined to automatically generate long video segment information corresponding to the historical video. Specifically, based on image association information and dialogue association information, the positional information of the target historical short video within a preset long video is determined, thereby extracting the preset long video to provide basic materials for the production of the new short video. Then, combined with video analysis technology, the selected long video segments are processed and synthesized to generate a mixed-cut short video containing a complete and engaging storyline with a fast pacing, suitable for promotional distribution to users.
[0073] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0074] Reference Figure 4 The diagram illustrates a structural block diagram of a video generation device embodiment according to the present invention, which may specifically include the following modules:
[0075] The association information acquisition module 401 is used to acquire target historical short videos and acquire association information between the target historical short videos and long video materials corresponding to preset long videos; the association information includes scene association information and dialogue association information.
[0076] The location information acquisition module 402 is used to generate location information of the target historical short video in the preset long video according to the screen association information and the dialogue association information; the location information is used to indicate the video range of at least one long video material.
[0077] The video generation module 403 is used to extract the preset long video based on the location information and generate a new video based on the extracted video segment.
[0078] In one embodiment of the present invention, the associated information acquisition module 401 may include the following sub-modules:
[0079] The image association information acquisition submodule is used to locate the first long video material range corresponding to the images contained in the target historical short video from the preset video site, and obtain the image association information.
[0080] In one embodiment of the present invention, the image association information acquisition submodule may include the following units:
[0081] The process involves: acquiring the image features of the target historical short video and acquiring a reference image feature library for the preset long video; performing a similarity search between the image features and the reference image feature library, determining that the target long video material corresponding to the target reference image feature with a similarity exceeding a preset first threshold is associated with the image corresponding to the image feature; acquiring the video time of the target long video material, determining a first long video material interval in the preset long video based on the video time of the target long video material, and obtaining image association information between the first long video material interval and the image corresponding to the image feature.
[0082] In one embodiment of the present invention, the associated information acquisition module 401 may include the following sub-modules:
[0083] The dialogue association information acquisition submodule is used to locate the second long video material range corresponding to the dialogue contained in the target historical short video from the preset video site, and obtain the dialogue association information.
[0084] In one embodiment of the present invention, the dialogue association information acquisition submodule may include the following units:
[0085] The dialogue association information acquisition unit is used to perform dialogue analysis on the target historical short video and the preset long video respectively, to obtain the first dialogue information of each line of dialogue in the target historical short video and the second dialogue information of each line of dialogue in the preset long video; to perform similarity calculation on the first dialogue information and the second dialogue information; if the similarity between the first dialogue information and the second dialogue information exceeds a preset second threshold, it is determined that the time interval in which the first dialogue information exists is associated with the dialogue content in the time interval in which the second dialogue information exists; to obtain the video time of the associated dialogue content, and to determine a second long video material interval in the preset long video based on the video time of the associated dialogue content, so as to obtain the dialogue association information between the second long video material interval and the associated dialogue content.
[0086] In one embodiment of the present invention, the associated information includes image association information and dialogue association information, wherein the image association information is used to indicate a first long video material range corresponding to the target historical short video image, and the dialogue association information is used to indicate a second long video material range corresponding to the target historical short video dialogue.
[0087] The location information acquisition module 402 may include the following sub-modules:
[0088] The location information generation submodule is used to merge the first long video material interval and the second long video material interval to obtain a union time interval; and to use the union time interval as the location information of the target historical short video in the preset long video.
[0089] In one embodiment of the present invention, the video interval is determined based on the video start time and video end time; the video generation module 403 may include the following sub-modules:
[0090] The video generation submodule is used to extract multiple video segments from the preset long video based on the video start time and the video end time; and to splice the multiple video segments to obtain the generated video.
[0091] In this embodiment of the invention, the video generation device provides that, by acquiring the association information between the target historical short video and the corresponding long video material of the preset long video, determines the position information of the target historical short video in the preset long video based on the aforementioned image association information and dialogue association information. The position information can be used to indicate the video range of at least one long video material. At this time, the preset long video can be cropped based on the generated position information to determine the video range of the original long video corresponding to the historical short video on the site, thereby acquiring the complete video plot. A new video is generated based on the cropped video segment to complete the plot information of the historical short video on the site, thereby realizing the processing of the historical short video on the site for promotion and distribution to users.
[0092] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0093] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0094] Memory 503 is used to store computer programs;
[0095] When processor 501 executes the program stored in memory 503, it performs the following steps:
[0096] Obtain target historical short videos, and obtain the association information between the target historical short videos and the long video materials corresponding to preset long videos; the association information includes image association information and dialogue association information;
[0097] Based on the image association information and the dialogue association information, the location information of the target historical short video in the preset long video is generated; the location information is used to indicate the video range of at least one long video material.
[0098] The preset long video is cropped based on the location information, and a new video is generated based on the cropped video segment.
[0099] In this embodiment of the invention, the electronic device provided by the present invention obtains the association information between the target historical short video and the long video material corresponding to the preset long video. Based on the aforementioned image association information and dialogue association information, it determines the position information of the target historical short video in the preset long video. The position information can be used to indicate the video range of at least one long video material. At this time, the preset long video can be cropped based on the generated position information to determine the video range of the original long video corresponding to the historical short video on the site, thereby obtaining the complete video plot. Based on the cropped video segment, a new video is generated to complete the plot information of the historical short video on the site, thereby realizing the processing of the historical short video on the site for promotion and distribution to users.
[0100] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0101] The communication interface is used for communication between the aforementioned terminal and other devices.
[0102] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0103] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0104] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the video generation methods described in the above embodiments.
[0105] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the video generation methods described in the above embodiments.
[0106] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0107] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0108] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0109] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A video generation method, characterized in that, The method includes: Obtain target historical short videos, and obtain the association information between the target historical short videos and the corresponding long video materials of preset long videos; the association information includes image association information and dialogue association information; wherein, the association information is used to indicate the degree of association between the target historical short videos and the preset long videos, and the degree of association is determined by a combination of image association degree and dialogue association degree; Based on the consistency verification and fusion of the image association information and the dialogue association information, the position information of the target historical short video in the preset long video is generated; the position information is used to indicate the video range of at least one long video material. The preset long video is cropped based on the location information, and a new video is generated based on the cropped video segment.
2. The method according to claim 1, characterized in that, The step of obtaining the association information between the target historical short video and the long video material corresponding to the preset long video includes: From the preset video site, the first long video material range corresponding to the images contained in the target historical short video is located to obtain the image association information.
3. The method according to claim 2, characterized in that, The step of locating the first long video material range corresponding to the images contained in the target historical short video to obtain image association information includes: Obtain the image features of the target historical short video, and obtain a reference image feature library for the preset long video; The image features are compared with the reference image feature library for similarity retrieval, and the target long video material corresponding to the target reference image feature whose similarity exceeds a preset first threshold is associated with the image corresponding to the image feature. The video time of the target long video material is obtained, and a first long video material interval is determined in the preset long video based on the video time of the target long video material. The image association information of the first long video material interval and the image corresponding to the image feature is obtained.
4. The method according to claim 1 or 2, characterized in that, The step of obtaining the association information between the target historical short video and the long video material corresponding to the preset long video includes: From the preset video site, the second long video material range corresponding to the lines contained in the target historical short video is located to obtain the line association information.
5. The method according to claim 4, characterized in that, The step of locating the second long video material range corresponding to the lines contained in the target historical short video to obtain line association information includes: The dialogue analysis is performed on the target historical short video and the preset long video respectively to obtain the first dialogue information of each line of dialogue in the target historical short video and the second dialogue information of each line of dialogue in the preset long video. Perform a similarity calculation on the first dialogue information and the second dialogue information; If the similarity between the first dialogue information and the second dialogue information exceeds a preset second threshold, then it is determined that the time interval in which the first dialogue information exists is related to the dialogue content in the time interval in which the second dialogue information exists. Obtain the video moment of the associated dialogue content, and determine the second long video material interval in the preset long video based on the video moment of the associated dialogue content, so as to obtain the dialogue association information between the second long video material interval and the associated dialogue content.
6. The method according to claim 1, characterized in that, The image association information is used to indicate the first long video material range corresponding to the target historical short video image, and the dialogue association information is used to indicate the second long video material range corresponding to the target historical short video dialogue. The step of generating the location information of the target historical short video within the preset long video based on the image association information and dialogue association information includes: The first long video clip interval and the second long video clip interval are merged to obtain the union time interval; The union time interval is used as the location information of the target historical short video in the preset long video.
7. The method according to claim 1 or 6, characterized in that, The video interval is determined based on the video start time and video end time; the step of trimming the preset long video according to the location information and generating a new video based on the trimmed video segment includes: Based on the video start time and the video end time, the preset long video is trimmed to obtain multiple trimmed video segments; The multiple video segments are spliced together to obtain the generated video.
8. A video generation apparatus, characterized in that, The device includes: The association information acquisition module is used to acquire target historical short videos and acquire association information between the target historical short videos and the long video materials corresponding to preset long videos; the association information includes screen association information and dialogue association information; wherein, the association information is used to indicate the degree of association between the target historical short videos and the preset long videos, and the degree of association is determined by a combination of screen association degree and dialogue association degree; The location information acquisition module is used to perform consistency verification and fusion based on the image association information and the dialogue association information to generate the location information of the target historical short video in the preset long video; the location information is used to indicate the video range of at least one long video material. The video generation module is used to extract segments from the preset long video based on the location information, and generate new videos based on the extracted video segments.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the video generation method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the video generation method as described in any one of claims 1-7.
11. A computer program, characterized in that, When executed by a processor, the computer program implements the video generation method as described in any one of claims 1-7.