Short video production method, electronic device and computer program product

By matching the main video to generate short imitation videos with logical similarity and strong sense of balance, the problems of high labor costs and low efficiency are solved, and high-efficiency production of high-quality short videos are achieved efficiently and automatically mass-produced, improving publicity effect.

CN119316631BActive Publication Date: 2025-09-02BEIJING YOUKU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411418778.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-09-02
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

The method of producing short videos in the prior art is high in labor costs and low efficiency, making it difficult to meet the needs of video platforms.

Method used

By obtaining target short videos with high user attention, matching them with the main video, extracting and generating short videos that are similar in logic, strong sense of compilation and matching users' preferences, using generative artificial intelligence models to predict emotions and add background music, we can achieve automatic mass production of high-quality short videos.

Benefits of technology

It realizes automatic mass production of high-quality short videos at low cost and high efficiency, improves publicity effects, and meets the publicity needs of video platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316631B_ABST
    Figure CN119316631B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a short video production method, electronic device, and computer program product. The method includes obtaining multiple target short videos, wherein the target short videos include short videos whose user attention level is greater than a specified threshold; matching each target short video in the multiple target short videos with a main video to obtain a matching result for each target short video; for any target short video in the multiple target short videos, when the target short video matches the main video, determining a first imitation short video of the target short video based on segment information of a video segment in the main video that matches the target short video; and generating multiple second imitation short videos based on multiple first imitation short videos of the multiple target short videos that match the main video. Thus, it is possible to achieve low-cost and high-efficiency automatic batch production of high-quality short videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a short video production method, electronic equipment, and computer program product. Background Art

[0002] With the rapid development of the short video industry, various short videos have greatly occupied people's fragmented time. Therefore, it has become a popular trend to use long videos such as TV dramas and variety shows to produce short videos, such as for the promotion of long videos.

[0003] The current method of producing short videos is usually to manually produce them based on the highlights of the main content of long videos such as TV series and variety shows. The labor cost is high, the efficiency is low, and the output is small, which makes it difficult to meet the needs of video platforms for short video production. Summary of the Invention

[0004] In view of this, the present disclosure proposes a short video production method, electronic device and computer program product, which can realize automatic batch production of high-quality short videos at low cost and high efficiency.

[0005] According to one aspect of the present disclosure, a short video production method is provided, comprising: obtaining a plurality of target short videos, wherein the target short videos include short videos of which the user attention level is greater than a specified threshold; matching each target short video in the plurality of target short videos with a main video to obtain a matching result for each target short video; wherein the matching result comprises: whether the target short video matches the main video, and, when the target short video matches the main video, segment information of a video segment in the main video that matches the target short video; for any target short video in the plurality of target short videos, when the target short video matches the main video, determining a first imitation short video of the target short video according to the segment information of the video segment in the main video that matches the target short video; and generating a plurality of second imitation short videos according to the plurality of first imitation short videos of the plurality of target short videos that match the main video.

[0006] In a possible implementation, the clip information includes the main film lines in the video clip, as well as the start time and the end time of the video clip; the matching of each target short video in the multiple target short videos with the main film video to obtain the matching result of each target short video includes: for any target short video in the multiple target short videos, extracting short video lines from the target short video to obtain a first line set corresponding to the target short video, the first line set including multiple short video lines appearing in the target short video; matching each short video line in the first line set with the main film lines in the main film line library; The character string is matched to obtain a line matching result of each short video line; wherein the main film line library includes all the main film lines in the main film video and the start time and end time of the video clip containing each main film line in all the main film lines; the line matching result includes: whether it matches the main film lines in the main film line library, and, when matching the main film lines in the main film line library, the start time and end time of the video clip containing the matched main film lines; the matching result of the target short video is determined based on the line matching result of each short video line in the first line set.

[0007] In one possible implementation, the segment information includes a main video frame in the video segment, and the start and end times of the video segment in which the main video frame is located. Matching each target short video in the multiple target short videos with the main video to obtain a matching result for each target short video includes: for any target short video in the multiple target short videos, if no short video lines are extracted from the target short video, or if the number of short video lines extracted from the target short video is less than a specified number of lines, performing similarity matching between the short video frames in the target short video and the main video frames in the main frame library to obtain a frame matching result for the short video frames in the target short video; wherein the main frame library includes all main video frames extracted at intervals from the main video; the frame matching result includes: whether the frame matches the main video frame in the main frame library, and, if the frame matches the main video frame in the main frame library, the frame time of the matched main video frame; and determining the matching result of the target short video based on the frame matching result of the short video frames in the target short video.

[0008] In a possible implementation, the feature video includes multiple video episodes, and the method of performing string matching on each short video line in the first line set with the feature film lines in the feature film line library to obtain a line matching result for each short video line includes: for any short video line in the first line set, performing string matching on the short video lines with the feature film lines in the feature film line library to obtain multiple feature film lines that match the short video lines, and then selecting, based on a target set video where the context lines of the short video lines are located, feature film lines that appear in the target set video from the multiple feature film lines that match the short video lines; and determining the line matching result for the short video lines based on the start time and end time of the video clip where the selected feature film lines that appear in the target set video are located.

[0009] In a possible implementation, the performing string matching of each short video line in the first line set with the main film lines in the main film line library to obtain the line matching results of each short video line includes: for any short video line in the first line set, performing string matching of the short video lines with the main film lines in the main film line library to obtain multiple main film lines that match the short video lines, performing similarity matching of the main film video frames containing the multiple main film lines that match the short video lines with the short video frames containing the short video lines to obtain the main film video frames with the highest similarity to the short video frames containing the short video lines; and determining the line matching results of the short video lines based on the main film lines corresponding to the main film video frame with the highest similarity to the short video frame containing the short video lines and the start time and end time of the video clip containing the short video lines.

[0010] In one possible implementation, the extracting of short video lines from a target short video to obtain a first line set corresponding to the target short video includes: extracting short video frames from the target short video at n frames per second to obtain multiple extracted short video frames, and performing character recognition on each of the short video frames extracted from the target short video to obtain recognized character strings and positions of the character strings in the multiple short video frames, where n is a positive integer; clustering and grouping the recognized character strings in the multiple short video frames according to the positions of the recognized character strings in the multiple short video frames to obtain multiple character string groups; and determining a first character string group containing the largest number of character strings among the multiple character string groups as the first line set.

[0011] In a possible implementation, the first imitation short video of the target short video is determined based on the segment information of the video segment that matches the target short video in the main video, including: when there are multiple video segments that match the target short video in the main video, according to the segment information of the multiple video segments that match the target short video in the main video, the multiple video segments that match the target short video in the main video are spliced ​​to obtain an initial imitation short video corresponding to the target short video; according to the top title set and / or flower character set of the target short video, the top title and / or flower character set are added to the initial imitation short video corresponding to the target short video to obtain the first imitation short video of the target short video.

[0012] In one possible implementation, the method further includes: determining a second string group among the multiple string groups, whose string positions are at the top of the short video frame and whose string repetition times exceed a specified threshold, as a top title set of the target short video; and / or determining string groups other than the first string group and the second string group among the multiple string groups as a flower character set in the target short video.

[0013] In a possible implementation, the first imitation short video of the target short video is determined based on the segment information of the video segment that matches the target short video in the main video, including: when there are multiple video segments that match the target short video in the main video, the multiple video segments that match the target short video in the main video are spliced ​​according to the segment information of the multiple video segments that match the target short video in the main video to obtain an initial imitation short video corresponding to the target short video; using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video according to the lines corresponding to the initial imitation short video and the auxiliary information of the main video; the auxiliary information includes a plot introduction and / or a plot commentary; determining the first background music that matches the emotion expressed in the initial imitation short video from a preset background music library, and adding the first background music to the initial imitation short video to obtain the first imitation short video of the target short video.

[0014] In one possible implementation, the generating of multiple second imitation short videos based on multiple first imitation short videos of multiple target short videos matching the main video includes: when more than a specified number of multiple first imitation short videos of multiple target short videos matching the main video are obtained, generating multiple second line combinations based on the first line combinations in the multiple first imitation short videos and auxiliary information related to the main video using a generative artificial intelligence model; the auxiliary information includes plot introduction and / or plot commentary; determining the video segment in which each line in each second line combination in the multiple second line combinations is located in the main video, and determining the multiple second imitation short videos corresponding to the multiple second line combinations based on the video segment in which each line in each second line combination in the multiple second line combinations is located in the main video.

[0015] In a possible implementation, determining the video segment in the feature video where each line in each of the multiple second line combinations is located includes: for any second line combination in the multiple second line combinations, performing string matching on each line in the second line combination with the feature lines in the feature line library to obtain a line matching result for each line in the second line combination, the line matching result including: whether it matches the feature lines in the feature line library, and, when matching the feature lines in the feature line library, the start time and end time of the matched feature lines and the video segment where the matched feature lines are located; and determining the video segment in the feature video where each line in the second line combination is located based on the line matching result of each line in the second line combination.

[0016] In a possible implementation, the method of determining a plurality of second imitation short videos corresponding to the plurality of second line combinations based on the video segments in the feature video where each line in each second line combination in the plurality of second line combinations is located includes: for any second line combination in the plurality of second line combinations, when the second line combination includes multiple lines, splicing the video segments in the feature video where each line in the second line combination is located to obtain an initial imitation short video corresponding to the second line combination; and adding a top title and / or a flower character set to the initial imitation short video corresponding to the second line combination based on the top title set and / or flower character set of a plurality of target short videos matching the feature video to obtain a second imitation short video corresponding to the second line combination.

[0017] In a possible implementation, the method of determining a plurality of second imitation short videos corresponding to the plurality of second line combinations based on the video segments in the feature video where each line in each second line combination in the plurality of second line combinations is located includes: for any second line combination among the plurality of second line combinations, when the second line combination includes multiple lines, splicing the video segments in the feature video where each line in the second line combination is located to obtain an initial imitation short video corresponding to the second line combination; using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video corresponding to the second line combination based on the line content of the second line combination and the auxiliary information; determining second background music from a preset background music library that matches the emotion expressed in the initial imitation short video corresponding to the second line combination, and adding the second background music to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video corresponding to the second line combination.

[0018] According to another aspect of the present disclosure, a short video production device is provided, including: an acquisition module for acquiring multiple target short videos, wherein the target short videos include short videos whose user attention level is greater than a specified threshold; a matching module for matching each target short video in the multiple target short videos with a main video to obtain a matching result for each target short video; wherein the matching result includes: whether the target short video matches the main video, and, when the target short video matches the main video, the segment information of the video segment in the main video that matches the target short video; a first imitation module for determining, for any target short video in the multiple target short videos, a first imitation short video of the target short video according to the segment information of the video segment in the main video that matches the target short video when the target short video matches the main video; a second imitation module for generating multiple second imitation short videos based on multiple first imitation short videos of the multiple target short videos that match the main video; the multiple first imitation short videos and the multiple second imitation short videos are used to promote the main video.

[0019] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0020] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0021] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0022] According to various aspects of the present disclosure, by utilizing a target short video whose user attention level is higher than a specified threshold and matches the main video, a high-quality first imitation short video is generated, which is similar in logic to the target short video, has a strong sense of rhythm, conforms to user preferences, etc., and matches the content of the main video. Then, multiple first imitation short videos that match the main video are used to generate more new second imitation short videos, which are similar in logic to the target short video, have a strong sense of rhythm, conform to user preferences, etc., and match the content of the main video. This achieves low-cost and high-efficiency automatic batch production of high-quality imitation short videos. In an exemplary application scenario, the promotional effect of using the first imitation short video and the second imitation short video to promote the main video can be improved, thereby meeting the promotional needs of the video platform for the main video.

[0023] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0025] Figure 1 A flowchart of a short video production method according to an embodiment of the present disclosure is shown.

[0026] Figure 2 A schematic diagram of a target short video according to an embodiment of the present disclosure is shown.

[0027] Figure 3 A schematic diagram illustrating a process of striping and OCR for a feature video according to an embodiment of the present disclosure is shown.

[0028] Figure 4 A schematic diagram showing a process of a short video production method according to an embodiment of the present disclosure is shown.

[0029] Figure 5 A block diagram of a short video production device according to an embodiment of the present disclosure is shown.

[0030] Figure 6 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0031] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0032] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0033] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C. In the description of this disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0034] It should be understood that the terms "first," "second," and the like in the claims, specification, and drawings of the present disclosure are used to distinguish between different objects, rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0035] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0036] The short video production method of the embodiment of the present disclosure can be deployed on various terminal devices through software or hardware modification. The terminal device involved in the embodiment of the present disclosure may refer to a device with a wireless connection function and / or a wired connection function. The wireless connection function refers to the ability to connect to other devices through wireless connection methods such as wifi and Bluetooth. The terminal device involved in the embodiment of the present disclosure can also communicate with other devices through a wired connection function. The terminal device involved in the embodiment of the present disclosure can be a touch screen, a non-touch screen, or a screenless device. The touch screen device can be controlled by clicking, sliding, etc. on the display screen with a finger, a stylus, etc. The non-touch screen device can be connected to an input device such as a mouse, keyboard, touch panel, etc., and the terminal device can be controlled through the input device. For example, a device without a screen can be a Bluetooth speaker without a screen. For example, the terminal device of the present application may include but is not limited to user equipment (UE), mobile device, user terminal, terminal, handheld device, tablet computer, laptop computer, PDA, computing device, etc.

[0037] The short video production method of the embodiment of the present disclosure can also be deployed on a server, which can be located in the cloud or locally, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., with a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. It can refer to a device with a wireless connection function, and the function of wireless connection means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server involved in the embodiment of the present disclosure can also have the function of communicating via a wired connection. For example, the server of the embodiment of the present disclosure can be located in the cloud, communicate with the terminal device, receive the target short video sent by the terminal device, and use the production method deployed on the server to generate multiple first imitation short videos and multiple second imitation short videos based on the target short video, and return them to the terminal device to display the generated multiple first imitation short videos and multiple second imitation short videos to the user in the terminal device, and also facilitate the user to use the generated multiple first imitation short videos and multiple second imitation short videos for application scenarios such as promoting feature videos.

[0038] The term "feature video" (or "long video") in this application refers to videos that are relatively long (e.g., over 10 minutes) with continuous, complete content, including but not limited to film and television dramas, variety shows, and performance recordings. The term "short video" in this application refers to videos that are relatively short (e.g., under 15 minutes) and are excerpts or clips from the feature video. This application does not impose any restrictions on the specific length or content of short videos or feature videos.

[0039] The lines in this application may include dialogues, monologues, narrations and other forms, and may be presented in the video in the form of subtitles or audio, and may be identified by relevant recognition technologies.

[0040] Figure 1 The flowchart of the short video production method according to an embodiment of the present disclosure is shown. The method can be applied to the above-mentioned terminal device or server and other electronic devices, such as Figure 1 As shown, the method includes: steps S11 to S14.

[0041] In step S11, a plurality of target short videos are obtained, where the target short videos include short videos whose user attention level is greater than a specified threshold.

[0042] The degree of user attention can be measured individually or comprehensively by the number of likes, favorites, reposts, and playbacks, etc., which is not limited in the embodiments of the present disclosure. It should be understood that those skilled in the art can set the specific value of the specified threshold according to actual needs. For example, if the number of likes is used to measure the degree of user attention, the specified threshold can be set to 100,000. In other words, short videos with more than 100,000 likes can be obtained as target short videos. Such short videos can be understood as popular short videos with high popularity.

[0043] In practical applications, short videos with a user attention level greater than a specified threshold can be downloaded from various social platforms, short video platforms, and other network platforms as target short videos. The disclosed embodiment does not limit the method for obtaining the target short videos.

[0044] Optionally, the target short videos can be screened to exclude unnecessary target short videos, such as those irrelevant to the main video content or those of poor quality. The screened target short videos are used as the multiple target short videos acquired in step S11, thereby reducing the data processing load in the subsequent step S12. Screening methods include, but are not limited to, manual selection and screening based on a trained artificial intelligence model.

[0045] It should be understood that the target short videos that attract a high degree of user attention and are related to the content of the main video usually include the highlights, or wonderful moments, in the main video. Since such target short videos have been verified by the market and have multiple factors such as smooth logic, strong sense of rhythm, and compliance with the preferences of users of various platforms, the embodiment of the present disclosure uses a method for generating imitation short videos by analyzing target short videos, which can enable the imitation short videos to also include highlight moments and also have multiple factors such as smooth logic, strong sense of rhythm, and compliance with the preferences of users of various platforms, which is conducive to improving the production efficiency, content quality and publicity effect of imitation short videos.

[0046] In step S12, each target short video among the multiple target short videos is matched with the feature video to obtain a matching result for each target short video.

[0047] The matching result includes: whether the target short video matches the main video, and, if the target short video matches the main video, the segment information of multiple video segments in the main video that match the target short video. A match between the target short video and the main video can be understood as the presence of at least one video segment in the main video that matches the target short video, and a mismatch between the target short video and the main video can be understood as the absence of a video segment in the main video that matches the target short video.

[0048] It is understandable that the main video may contain lines, so the target short video can be matched with the main video based on the main line. The above-mentioned segment information may include the main line in the video segment, as well as the start time and end time of the video segment. In a possible implementation, the above-mentioned step S12 matches each target short video in the multiple target short videos with the main video to obtain a matching result for each target short video, which may include:

[0049] Step S121: extracting short video lines from any target short video among the multiple target short videos to obtain a first line set corresponding to the target short video, where the first line set includes multiple short video lines appearing in the target short video;

[0050] Step S122: Perform string matching on each short video line in the first line set with the main film lines in the main film line library to obtain a line matching result for each short video line; wherein the main film line library includes all the main film lines in the main film video and the start and end times of the video clip containing each main film line in all the main film lines; the line matching result includes: whether the main film lines in the main film line library match, and, if the main film lines in the main film line library match, the start and end times of the video clip containing the matched main film lines;

[0051] Step S123 : determining a matching result of the target short video based on the line matching results of each short video line in the first line set.

[0052] In step S121, text recognition technology known in the art, such as optical character recognition (OCR) technology, can be used to extract short video lines from the target short video. For example, frames can be extracted from the target short video and OCR technology can be used to perform string recognition on the extracted short video frames. By recognizing the subtitles in the short video frames, the short video lines appearing in the target short video can be obtained. It should be understood by those skilled in the art that short video lines can also be extracted by other methods, such as using speech recognition technology to recognize the speech content in the audio of the target short video to extract short video lines.

[0053] The extracted short video lines can be in the form of strings. Each video frame containing lines can correspond to one or more strings. Multiple video frames can correspond to the same string, that is, multiple video frames can correspond to the same line. Each sentence in natural language can be identified as one or more strings. Strings in the same row of a video frame can be identified as a line, or multiple lines of strings adjacent to each other on the vertical axis can be identified as a line. Furthermore, strings in the same row or multiple lines of strings can be divided into at least one line based on punctuation marks (such as commas, periods, and question marks) identified in the strings.

[0054] Considering that the subtitles of the lines in the short video may appear at any position in the short video, and the short video may also have top titles and fancy words, fancy words can be understood as text content in the short video with any color, font, font size, etc. except the lines and top titles, for example, Figure 2 A schematic diagram of a target short video is shown, such as Figure 2 As shown, the target short video has the top title "AA Variety Show's Funny Scenes", the short video line "Do you believe in light?", and the colorful characters "Hahaha". Therefore, all the text recognized from the target short video can be grouped by vertical coordinate aggregation. The group containing the largest amount of text content (such as the character string below) is regarded as the short video lines, the ones at the top of the short video and repeated more than a specified threshold (such as 5 times) are regarded as the top title of the target short video, and the other groups are regarded as colorful characters. Specifically, the above step S121 extracts the short video lines from the target short video to obtain the first line set corresponding to the target short video, including:

[0055] Step S1211: extracting short video frames from the target short video at a rate of n frames per second to obtain multiple extracted short video frames, and performing character recognition on each of the short video frames extracted from the target short video to obtain recognized character strings and positions of the character strings in the multiple short video frames, where n is a positive integer;

[0056] Step S1212, clustering and grouping the character strings recognized in the multiple short video frames according to the positions of the character strings recognized in the multiple short video frames to obtain multiple character string groups;

[0057] Step S1213, determining the first character string group containing the largest number of character strings among the multiple character string groups as the first line set;

[0058] And, the method may further include:

[0059] Step S1214: determining a second string group in which the string position is at the top of the short video frame and the number of repetitions of the string exceeds a specified threshold as the top title set of the target short video; and / or,

[0060] Step S1215 : Determine the character string groups other than the first character string group and the second character string group in the multiple character string groups as a set of fancy characters in the target short video.

[0061] In step S1211, those skilled in the art can set a specific value of n based on historical experience. For example, n can be set to 2, that is, short video frames can be extracted from the target short video at a rate of two frames per second. For a 10-second target short video, 20 short video frames can be extracted. Then, OCR technology can be used to perform character recognition on each short video frame extracted from the target short video to obtain the character strings recognized in each short video frame and the position of the character strings. The character strings are then clustered and grouped based on their positions (that is, all the character strings in the target short video are grouped by vertical coordinate aggregation). It should be understood that the positions of the various character strings in the same character string group are the same or similar in the vertical coordinate direction, and the positions of the character strings in different character string groups are different.

[0062] Since the target short video has the largest number of lines, the strings in the first string group containing the largest number of strings can be determined as the short video lines in the target short video; since the top title of a short video is usually at the top of the short video and is repeated multiple times (i.e., appears in multiple short video frames), the strings in the second string group whose positions are at the top of the short video frame and whose repetition times exceed a specified threshold can be determined as the top title in the target short video; since the position and number of flower characters in the target short video are not fixed, the strings in string groups other than the first and second string groups can be determined as the flower characters in the target short video. In this way, the short video lines, top titles, and flower characters in the target short video can be accurately identified.

[0063] It should be understood that if there is no top title in the target short video, the character string groups other than the first character string group can be determined as flowery characters. If there is no flowery characters in the target short video, the short video lines in the first character group and the top title in the second character string group can be obtained. If there is no top title and flowery characters in the target short video, a character string group may be clustered, and the clustered character string group is also the first line set.

[0064] In step S122, the main film dialogue library can pre-identify and split all the main film dialogues in the main film video through video stripping or OCR technology to obtain all the main film dialogues and the start and end times of the video segments where each main film dialogue is located, wherein a video segment contains the same main film dialogue. The video frame displaying the subtitles of the same main film dialogue may be split into multiple video segments. For example, when "I haven't eaten yet, let's eat together" is displayed as a subtitle on multiple consecutive video frames, these multiple consecutive video frames may be split into two video segments based on voice features. The first video segment corresponds to the voice "I haven't eaten yet", and the second video segment corresponds to the voice "Let's eat together". The way of splitting the video segments is related to the specific way of video stripping, and the present disclosure does not impose any restrictions on this.

[0065] Since the main film lines obtained by the video stripping technology have a certain error rate, OCR technology can be used to correct the main film lines obtained by the video stripping technology to ensure the accuracy of the identified main film lines, thereby improving the success rate of subsequent matching of target short videos. Figure 3 A process of stripping and OCR for positive video is shown as follows: Figure 3 As shown, the episode list of the main video can be obtained and stored on the local device, and then multiple episodes of the main video can be downloaded to the local device in batches and the videos being downloaded and downloaded successfully are marked (that is, the videos being downloaded and downloaded successfully are marked) and stored on the local device. At the same time, the multiple episodes downloaded to the local device are stripped or OCR recognition is performed, and the video clips stripped are stored in a storage server (such as a server using an object storage service (OSS)), and the OCR recognition results are used to correct the lines of the video stripping results, and the corrected main film lines are stored in the local database; wherein, multiple devices can be used to concurrently perform video stripping and OCR recognition, and data can be synchronized between multiple devices to avoid duplication or omission.

[0066] In actual applications, each line of the main film identified by video stripping or OCR technology and the corresponding start and end time (i.e., the start and end time of the video clip where the main film lines are located) can be written into the database in the form of a JSON array to obtain a main film line library. For example, the main film lines and the corresponding start and end times stored in the form of a JSON array can be expressed as: [{"text":"Have you eaten?","startTime":102400,"endTime":105400},{"text":"I haven't eaten yet,","startTime":105400,"endTime":108880},{"text":"Where do you want to go for dinner?","startTime":108880,"endTime":110280}], where text represents the main film lines, startTime represents the start time, and endTime represents the end time.

[0067] In actual applications, after obtaining the JSON arrays corresponding to all the main lines in the main video, the JSON arrays corresponding to each main line can be combined to use video processing technologies known in the art, such as FFmpeg (an open source computer program for processing video and audio) software, to divide the main video into video segments, and the start and end time of each video segment is the above-mentioned startTime and endTime, and each segmented video segment can be stored in a storage server (such as a server using Object Storage Service (OSS)) to facilitate subsequent splicing of video segments to obtain simulated short videos.

[0068] Based on the above-mentioned main film dialogue library, in step S122, any string matching technique known in the art can be used to perform string matching between each short video dialogue in the first dialogue set and the main film dialogue in the main film dialogue library to obtain a dialogue matching result for each short video dialogue. For example, the similarity between each short video dialogue in the first dialogue set and the main film dialogue in the main film dialogue library can be calculated, and the main film dialogue whose similarity exceeds a specified similarity threshold (e.g., 80%) is determined to be a main film dialogue that matches the short video dialogue. That is, if the similarity between a short video dialogue and a main film dialogue exceeds 80%, the short video dialogue is considered to match the main film dialogue (i.e., the match is successful); otherwise, the short video dialogue is considered to not match the main film dialogue. The dialogue corresponding to a video clip of the main film can be considered as a "line of dialogue." The division of short video dialogue into "sentences" can be aligned with the principles used for striping the main film, ensuring that a "sentence" in the short video dialogue corresponds as closely as possible to a "sentence" in a video clip of the main film, thereby improving matching accuracy. For example, both a short video dialogue and a main film dialogue can be divided into sentences demarcated by punctuation marks in natural language. For example, when extracting short video dialogue using subtitles from a short video using techniques such as OCR, strings ending with specific punctuation marks (such as commas, greetings, and periods) can be identified as "sentences." When striping the main film based on speech, these punctuation marks form sentence breaks and are therefore also segmented into video clips, each corresponding to a line from the main film. For example, if "Have you eaten?", "I haven't eaten yet," and "Where do you want to eat?" are three lines from the main film, then if a short video dialogue is identified as "Have you eaten?", "Have you eaten?" is matched against "Have you eaten?", "I haven't eaten yet," and "Where do you want to eat?", respectively. If the subtitles of a short video are "I haven't eaten yet, where do you want to go for dinner?", it can be identified as "I haven't eaten yet" and "Where do you want to go for dinner?" and matched with "Have you eaten?", "I haven't eaten yet," and "Where do you want to go for dinner?" respectively.

[0069] Considering that feature videos are usually long and may include multiple episodes, the same feature line may appear multiple times in different episodes of the feature video, so the same short video line may match multiple feature lines. Therefore, in one possible implementation, in step S122, string matching is performed on each short video line in the first line set with the feature lines in the feature line library to obtain the line matching results for each short video line, which may include:

[0070] For any short video line in the first line set, after matching the short video line with the main film lines in the main film line library to obtain multiple main film lines that match the short video line, the main film lines that appear in the target set video are selected from the multiple main film lines that match the short video line based on the target set video containing the context lines of the short video line. The line matching result of the short video line is determined based on the start and end times of the video clip containing the selected main film lines that appear in the target set video. In this way, the main film lines that match the overall content of the target short video can be selected from the multiple main film lines that match the short video line.

[0071] Among them, the context lines of the short video lines can be understood as the adjacent lines that appear above and below the short video lines in the target short video. It can be understood that the short video lines and the context lines of the short video lines usually appear in the same video episode or the same video segment. Therefore, the target set video where the context lines appear can be used to select the main film lines that appear in the target set from the multiple main film lines that match the short video lines. In this way, the main film lines that match the short video lines and the starting time and ending time of the video segment where the main film lines are located can be obtained.

[0072] Among them, the video clip where the context lines are located can be obtained by performing string matching between the context lines and the main film lines, and then the target set video to which the video clip where the context lines are located belongs can be obtained. It should be understood that if the context lines only appear in a certain episode of video, then the episode where the context lines appear is also the target set video. If the context lines also appear in multiple episodes of video, then the episode where the context lines appear the most times can be used as the target set video.

[0073] Considering that the same short video line may match multiple lines from the main film, and multiple lines from the main film may also appear in the same episode or the same video segment, this means that the same line from the main film may appear in the same episode or the same video segment. In order to more accurately select the main film lines that match the target short video content from the multiple lines that match the short video lines, in one possible implementation, step S122 performs string matching on each short video line in the first line set with the main film lines in the main film line library to obtain the line matching results for each short video line, which may include:

[0074] For any short video line in the first line set, after string matching the short video line with the main film lines in the main film line library to obtain multiple main film lines that match the short video line, similarity matching is performed on the main film video frames containing each of the multiple main film lines that match the short video line with the short video frame containing the short video line, obtaining the main film video frame with the highest similarity to the short video frame containing the short video line. The line matching result of the short video line is determined based on the main film lines corresponding to the main film video frame with the highest similarity to the short video frame containing the short video line and the start and end times of the video clip containing the short video line. In this way, the main film lines that match the overall content of the target short video can be accurately selected from the multiple main film lines that match the short video lines.

[0075] It is understandable that although the same short video line may match multiple lines of the main film, the main film video frames when the main film lines appear are usually different. Therefore, by calculating the similarity between the main film video frames containing the multiple lines that match the short video lines and the short video frames containing the short video lines, the main film lines corresponding to the main film video frames with the highest similarity can be selected from the matched multiple lines of the main film as the main film lines that correctly match the short video lines. Then, based on the main film lines corresponding to the main film video frames with the highest similarity and the start and end times of the video clip containing the main film lines, the line matching results of the short video lines are obtained. Among them, those skilled in the art can use any similarity calculation method known in the art to calculate the similarity between the main film video frames and the short video frames, and the embodiments of the present disclosure are not limited to this.

[0076] For example, if a line in a short video matches the lines A1, A2 and A3 in the main film, and the lines A1, A2 and A3 in the main film correspond to three segments B1, B2 and B3, respectively, each containing 20 frames of the main film video, and the short video corresponding to the short video line contains 20 frames of short video frames, the similarity between the first frame of the short video and the 20 frames of the three segments B1, B2 and B3 can be calculated, and the three segments B1, B2 and B3 are selected respectively. 1. The positive video frame with the highest similarity to the first short video frame among the 20 positive video frames of B1, B2, and B3 (such as the first positive video frame in B1, the second positive video frame in B2, and the third positive video frame in B3) is used as the positive video frame that matches the first short video frame in the three positive segments B1, B2, and B3. Then calculate the similarity between the second short video frame and the 20 positive video frames of the three positive segments B1, B2, and B3 respectively. Frame similarity is determined by selecting the positive video frame with the highest similarity to the first short video frame from the 20 positive video frames of the three positive film segments B1, B2, and B3 (e.g., the second positive video frame in B1, the third positive video frame in B2, and the fourth positive video frame in B3) as the positive video frame that matches the second short video frame in the three positive film segments B1, B2, and B3. Similarly, the positive video frame that matches each short video frame in the three positive film segments B1, B2, and B3 can be obtained. Then, the average similarity corresponding to the positive video frames that match each short video frame in the three positive film segments B1, B2, and B3 can be calculated. The positive line (positive line A1) corresponding to the positive segment with the largest average similarity (e.g., positive segment B1) is taken as the positive line corresponding to the positive video frame with the highest similarity, that is, the positive line that correctly matches the short video line.

[0077] It should be understood that for each short video line in the first line set, the line matching result for each short video line can be obtained according to the implementation method of the above step S122. Therefore, in step S123, the matching result for the target short video can be obtained based on the line matching results for each short video line in the first line set. And for each target short video obtained in step S11, the matching result for each target short video can be obtained according to the above steps S121 to S122.

[0078] It is understandable that there may be some lines in the target short video that do not belong to the main film. Therefore, there may be some short video lines in the first line set that do not match the main film lines in the main film line library. In this case, the short video lines that do not match the main film lines in the main film line library can be directly filtered out; and, since the multiple target short videos in step S11 are obtained based on the user's attention level, the video content of the target short video may also be completely mismatched with the main film, so that all the short video lines in the first line set do not match the main film lines in the main film line library. In this case, the target short video that does not match the main film video at all can be directly filtered out, that is, the target short video that does not match the main film video at all is not processed subsequently, so that the target short video that matches the main film video can be screened out from the multiple target short videos obtained in step S11.

[0079] Taking into account that some of the multiple target short videos obtained in step S11 may have no lines or very few lines, in these cases, video clips that match the highlight moments in the target short videos can be found by comparing video frames. For example, each short video frame in the target short video can be compared with all the main video frames extracted from the main video, and the main video frames with a similarity higher than 80% are used as the conditions for successful matching, so that the first imitation short video can be produced without lines. Specifically, in the above step S12, the clip information can also include the main video frame in the video clip, and the start time and end time of the video clip where the main video frame is located. The above matching of each target short video in the multiple target short videos with the main video to obtain the matching results of each target short video can include:

[0080] Step S124: For any target short video among the multiple target short videos, if no short video lines are extracted from the target short video, or if the number of short video lines extracted from the target short video is less than a specified number of lines, similarity matching is performed between the short video frames in the target short video and the main video frames in the main film frame library to obtain a frame matching result of the short video frames in the target short video; wherein the main film frame library includes all main film video frames extracted at intervals from the main film video; the frame matching result includes: whether the frame matches the main film video frame in the main film frame library, and, if the frame matches the main film video frame in the main film frame library, the frame time of the matched main film video frame;

[0081] Step S125 , determining a matching result of the target short video according to the frame matching results of the short video frames in the target short video.

[0082] Among them, failure to extract short video lines from the target short video can be understood as failure to identify short video lines from the target short video, and the number of short video lines extracted from the target short video is less than the specified number of lines can be understood as the number of short video lines identified from the target short video is less than the specified number of lines. Among them, those skilled in the art can set a specific value for the specified number of lines based on actual experience. For example, it can be set to 5 sentences, and this is not limited in the embodiments of the present disclosure.

[0083] In actual applications, the main film video frames can be extracted from the entire main film video at intervals according to a certain frame extraction method (such as extracting m frames per second) and stored (for example, all the extracted main film video frames can be stored in a storage server) to obtain a main film frame library. The embodiment of the present disclosure does not limit the construction method of the main film frame library.

[0084] In step S124, a similarity calculation method known in the art can be used to perform similarity matching on each short video frame in the target short video with the main video frame in the main film frame library to obtain a frame matching result of each short video frame in the target short video. Alternatively, each short video frame in a plurality of short video frames extracted at intervals from the target short video can be performed similarity matching on the main video frame in the main film frame library to obtain a frame matching result of each short video frame extracted from the target short video. Specifically, the similarity between the short video frame in the target short video and the main video frame in the main film frame library can be calculated, and the similarity exceeding The main video frame with a specified similarity threshold (such as 80%) is determined to be the main video frame that matches the short video frame, that is, if the similarity between the short video frame and a main video frame exceeds 80%, the short video frame is considered to match the main video frame, otherwise, the short video frame is considered to not match the main video frame; it should be understood that if the short video frame matches a main video frame, the main video frame that matches the short video frame and the frame time of the main video frame can be obtained, and the frame time is the time when the main video frame appears in the main video, wherein, for a main video containing multiple episodes of video, the frame time may include a certain moment in a certain episode of the main video.

[0085] Among them, for each short video frame in the target short video or each extracted short video frame, the frame matching result of the short video frame in the target short video can be obtained according to the implementation method of the above-mentioned step S124, that is, whether it matches the main video frame in the main film frame library, and, when matching the main video frame in the main film frame library, the frame time of the matched main video frame; and then in step S125, based on the frame time of the main video frame that matches the short video frame in the target short video, the video segment indicated by the frame time of the main video frame that matches the short video frame in the target short video can be determined.

[0086] In one example, the indicated video segment may be a video segment between the maximum and minimum values ​​of all frame times of the matching main video frame. For example, if the frame times of the main video frame matching the short video frame in the target short video include: 10:24:00, 10:24:10, 10:24:20, and 10:24:30, then the video segment indicated by the frame times of the main video frame matching the short video frame in the target short video is the video segment indicated from 10:24:00 to 10:24:30, with the starting time of the video segment being 10:24:00 and the ending time being 10:24:30.

[0087] In another example, the video clips obtained by splitting the main video in the previous article can be used. The indicated video clips can be video clips in which frame moments fall, such as the video clips obtained by splitting the video strips in which the four moments fall, such as 10:24:00, 10:24:10, 10:24:20, and 10:24:30.

[0088] In addition, it should be understood that the target short video may contain at least one video segment in the main video, and the frame moments of the main video frames in the same video segment are adjacent (the frame moments are adjacent, for example, the duration between two adjacent frame moments may be less than or equal to the unit duration of the interval frame extraction). Therefore, the video segments corresponding to the multiple adjacent frame moments can be determined based on the frame moments of the main video frames, and the starting moment and ending moment of the video segment can also be obtained, so as to obtain at least one video segment in the main video that matches the target short video and the starting moment and ending moment of the video segment.

[0089] As mentioned above, the target short video may also contain some content that does not belong to the main video, that is, the target short video may also contain some short video frames that do not match the main video frames in the main film frame library. In this case, the short video frames that do not match the main video frames in the main film frame library can be directly filtered out; and, since the multiple target short videos in step S11 are obtained based on the user's attention level, the video content of the target short video may also be completely mismatched with the main video, so that all short video frames in the target short video do not match the main video frames in the main film frame library. In this case, the target short video that is completely mismatched with the main video can be directly filtered out, that is, the target short video that is completely mismatched with the main video will not be processed subsequently, so that the target short video that matches the main video can be screened out from the multiple target short videos obtained in step S11.

[0090] In step S13, for any target short video among the multiple target short videos, when the target short video matches the main video, a first imitation short video of the target short video is determined based on the segment information of the video segment matching the target short video in the main video.

[0091] As described above, there may be one or more video segments in the main video that match the target short video. Therefore, when there is only one video segment in the main video that matches the target short video, the corresponding video segment can be directly obtained from the storage server used to store each video segment in the main video according to the segment information of the video segment that matches the target short video in the main video, and the obtained video segment can be determined as the first imitation short video of the target short video; when there are multiple video segments in the main video that match the target short video, the corresponding multiple video segments can be obtained from the storage server according to the segment information of the multiple video segments that match the target short video in the main video, and the obtained multiple video segments can be spliced ​​to obtain the first imitation short video of the target short video.

[0092] Considering that short videos in actual situations usually contain top titles and / or fancy words, it is also possible to refer to the top title and / or fancy words in the target short video to generate a first imitation short video with the top title and / or fancy words added thereto, so as to obtain a first imitation short video with better promotional effect. Therefore, in a possible implementation, the above step S13, which determines the first imitation short video of the target short video based on the segment information of the video segment in the main video that matches the target short video, may include:

[0093] Step S131: If there are multiple video segments in the main video that match the target short video, the multiple video segments in the main video that match the target short video are spliced ​​together according to the segment information of the multiple video segments in the main video that match the target short video to obtain an initial imitation short video corresponding to the target short video;

[0094] Step S132: adding the top title and / or the fancy words to the initial imitation short video corresponding to the target short video according to the top title set and / or the fancy words set of the target short video, to obtain a first imitation short video of the target short video.

[0095] In step S131, based on the segment information of the multiple video segments that match the target short video in the main video, the multiple video segments that match the target short video in the main video are spliced ​​together. The multiple video segments that match the target short video can be spliced ​​together in the order in which the multiple video segments that match the target short video appear in the target short video, so as to facilitate obtaining an initial imitation short video that conforms to the content logic in the target short video. In practical applications, those skilled in the art can adopt any video processing technology known in the art. For example, FFmpeg software can be used to splice the multiple video segments that match the target short video in the main video, that is, all the matched video segments can be combined by FFmpeg software. This is not limited in the embodiment of the present disclosure.

[0096] As described above, the top title set and / or flower character set of the target short video can be obtained by grouping the character strings identified from the target short video. If the top title set contains top titles that are repeated multiple times, then in step S132, the top titles that are repeated multiple times in the top title set can be directly added to the top position of the initial imitation short video (the top position can be the position where the top title is located in the target short video) to obtain the first imitation short video with the top title added.

[0097] It can be seen that the flower characters in the target short video are usually associated with the short video lines in the target short video. Therefore, the flower character set can also record the association between the flower characters and the short video lines identified in the same short video frame. Then, in step S132, based on the association between the flower characters and the short video lines, the flower characters associated with the short video lines can be added to the video frame where the main film lines that match the short video lines are located in the initial imitation short video, so as to obtain a first imitation short video with the flower characters added. It should be understood that the first imitation short video can have only the top title added, only the flower characters added, or both the top title and the flower characters added.

[0098] In actual applications, some target short videos may not originally have top titles and / or fancy words. If the target short video does not have top titles and / or fancy words, an artificial intelligence model can be used to generate top titles and / or fancy words based on the lines content, plot introduction, plot comments and other information in the target short video, so that in step S132, the top titles and / or fancy words generated by the model are added to the initial imitation short video corresponding to the target short video to obtain the corresponding first imitation short video.

[0099] Considering that in actual situations, background music (BGM) is usually added to short videos to enhance the emotional expression of the short videos, therefore, in a possible implementation, in the above step S13, determining the first imitation short video of the target short video based on the segment information of the video segment in the main video that matches the target short video may include:

[0100] Step S133: If there are multiple video segments in the main video that match the target short video, the multiple video segments in the main video that match the target short video are spliced ​​together based on the segment information of the multiple video segments in the main video to obtain an initial imitation short video corresponding to the target short video. The implementation of step S133 can refer to the above-mentioned step S131 and is not described in detail here.

[0101] Step S134, using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video based on the lines corresponding to the initial imitation short video and auxiliary information of the main video; the auxiliary information includes a plot introduction and / or plot commentary;

[0102] Step S135 , determining a first background music that matches the emotion expressed in the initial imitation short video from a preset background music library, and adding the first background music to the initial imitation short video to obtain a first imitation short video of the target short video.

[0103] In step S134, those skilled in the art can use any known, pre-trained generative artificial intelligence model in the art to predict the emotions expressed in the initial imitation short video based on the corresponding lines of the initial imitation short video and the auxiliary information of the feature video. The embodiment of the present disclosure does not limit the type, structure, etc. of the generative artificial intelligence model. Among them, the generative artificial intelligence model can be instructed to predict the emotions expressed in the initial imitation short video by inputting a command prompt into the generative artificial intelligence model, and the embodiment of the present disclosure does not limit this. The artificial intelligence model can be pre-trained based on line samples, auxiliary information samples and emotion labels based on the training methods in the prior art.

[0104] In practical applications, plot descriptions and comments about the main film can be obtained from various online platforms, such as drama review platforms, social platforms, and video platforms. These comments can be those with a certain threshold of likes or reposts. These comments can be considered high-quality comments that represent the hot topics of most users and can express users' opinions on the highlights of the main film. This can be used to leverage auxiliary information from the main film to help the generative AI model more accurately predict the emotions expressed in the initial imitation short video based on the lines in the initial imitation short video.

[0105] Among them, each background music in the background music library can have a business identifier, and the business identifier can be used to describe the usage scenario of the background music, or to describe the short video with what emotion the background music is suitable for, that is, the business identifier indicates the emotion expressed by the background music. Therefore, in step S135, the emotion expressed by the initial imitation short video can be matched with the background music with the business identifier in the background music library to determine the first background music that matches the emotion expressed by the initial imitation short video from the background music library, and add the first background music to the initial imitation short video to obtain the first imitation short video of the target short video. For example, the background music library contains background music 1, background music 2, and background music 3. The business identifier of background music 1 is "suitable for scenes with sad emotions", the business identifier of background music 2 is "suitable for scenes with joy emotions", and the business identifier of background music 3 is "suitable for scenes with loneliness emotions". If the emotion expressed by the initial imitation short video is joy, background music 2 can be matched to match the emotion expressed by the initial imitation short video.

[0106] As described above, there may be one video segment in the main video that matches the target short video. Therefore, when there is only one video segment in the main video that matches the target short video, the one video segment that matches the target short video can be directly determined as the initial imitation short video, and then the above-mentioned step S132 and / or the above-mentioned steps S134 to S135 can be executed on the initial imitation short video to obtain the first imitation short video of the target short video.

[0107] It should be understood that for the initial imitation short video, step S132 can be executed simultaneously to add a top title and / or fancy words, and steps S134 to S135 can be executed to add a matching first background music. That is to say, the first imitation short video can include the first background music, as well as the top title and / or fancy words; of course, the first imitation short video can also only include the top title and / or fancy words, or only include the first background music, and this is not limited to the embodiments of the present disclosure.

[0108] In step S14, a plurality of second imitation short videos are generated based on a plurality of first imitation short videos of a plurality of target short videos that match the main video.

[0109] It should be understood that through the above steps S12 to S13, the first imitation short video of each target short video in the multiple target short videos matching the main video can be determined from the multiple target short videos obtained in step S11, that is, multiple first imitation short videos can be obtained; among them, the advantage of the first imitation short video generated based on the target short video is that the target short video has been verified by the market, and the quality of the first imitation short video obtained after imitation is also high, but the disadvantage is that one target short video can only generate one corresponding first imitation short video. In order to increase the output of short videos, after obtaining enough first imitation short videos of target short videos matching the main video, these first imitation short videos can be handed over to the artificial intelligence model for mixed editing (that is, fragments from different first imitation short videos are spliced), so that more imitated second imitation short videos can be obtained, which is conducive to significantly increasing the production quantity of short videos.

[0110] Therefore, in one possible implementation method, the above-mentioned generation of multiple second imitation short videos based on multiple first imitation short videos of multiple target short videos matching the main video may include: using an artificial intelligence model to perform video remixing on the multiple first imitation short videos to obtain multiple second imitation short videos, wherein those skilled in the art can adopt artificial intelligence technology known in the art to develop and train artificial intelligence models for video remixing, and this is not limited to the embodiments of the present disclosure.

[0111] Considering that the production efficiency of directly mixing and editing multiple first imitation short videos to obtain multiple second imitation short videos is low, the embodiment of the present disclosure also proposes that a new second imitation video can be generated by using a generative artificial intelligence model to generate a new second line combination based on the first line combination in the first imitation short video in combination with the plot introduction and plot commentary. Compared with using the model to process the video, the model has higher efficiency and lower cost in processing text such as lines, which is conducive to improving the generation efficiency of the second imitation short video and reducing the generation cost. Therefore, in a possible implementation method, the above-mentioned step S14, generating multiple second imitation short videos based on multiple first imitation short videos of multiple target short videos matching the main video, can include:

[0112] Step S141: When a plurality of first imitation short videos are obtained that match a specified number of target short videos with the main video, a plurality of second dialogue combinations are generated using a generative artificial intelligence model based on the first dialogue combinations in the plurality of first imitation short videos and auxiliary information related to the main video; the auxiliary information includes plot introductions and / or plot comments;

[0113] Step S142, determining the video clips in which each line in each of the multiple second line combinations is located in the main video, and determining multiple second imitation short videos corresponding to the multiple second line combinations based on the video clips in which each line in each of the multiple second line combinations is located in the main video.

[0114] In step S141, those skilled in the art may use any generative artificial intelligence model known in the art to generate multiple second dialogue combinations based on the first dialogue combinations in the multiple first imitation short videos and auxiliary information related to the feature video. The disclosed embodiment does not limit the type, structure, etc. of the generative artificial intelligence model. Specifically, the generative artificial intelligence model may be instructed to automatically generate multiple second dialogue combinations based on the first dialogue combinations in the multiple first imitation short videos and the drama review introduction and / or plot commentary related to the feature video by inputting a command prompt into the generative artificial intelligence model. The disclosed embodiment does not limit this.

[0115] In actual applications, when a command prompt is input into the generative artificial intelligence model to instruct it to generate multiple second line combinations, the input prompt may include, in addition to auxiliary information indicating the first line combinations of the multiple first imitation short videos and the main video, some constraint information so that the generative artificial intelligence model can output second line combinations that meet the actual imitation requirements. For example, the constraint information may include the degree of change of the second line combination relative to the first line combination and the number of second line combinations generated. For example, the prompt input into the generative artificial intelligence model may be "Please use the 100 input line combinations (i.e., the first line combinations) combined with the input plot introduction and plot commentary to generate 1,000 new line combinations (i.e., the second line combinations), where the degree of change of the new line combinations relative to the input line combinations does not exceed 10% (i.e., more than 10% of the input line combinations cannot be changed)." It should be understood that those skilled in the art can design the specific content of the prompt input into the generative artificial intelligence model according to actual needs, and the embodiments of the present disclosure are not limited to this.

[0116] In step S142, determining the video segment in the main film where each line of each second line combination in the multiple second line combinations is located may include: for any second line combination in the multiple second line combinations, performing string matching on each line in the second line combination with the main film lines in the main film line library to obtain a line matching result for each line in the second line combination, the line matching result including: whether it matches the main film lines in the main film line library, and, when matching the main film lines in the main film line library, the start time and end time of the matched main film lines and the video segment where the matched main film lines are located; based on the line matching result of each line in the second line combination, determining the video segment in the main film where each line in the second line combination is located.

[0117] The string matching method described in step S122 above can be used to perform string matching on each line in the second line combination with the main film lines in the main film line library to obtain a line matching result for each line in the second line combination. This will not be described in detail here. Furthermore, after obtaining the line matching result for each line in the second line combination, the corresponding video segment can be retrieved from a storage server for storing video segments based on the start and end times of the video segment containing the successfully matched main film line. In other words, based on the line matching result for each line in the second line combination, the video segment in the main film video containing each line in the second line combination can be determined.

[0118] It should be understood that the second line combination is generated based on the first line combination, and the first line combination contains the lines of the main film, wherein the text in the second line combination can be understood as a rearranged combination of the original text in the first line combination, similar to a line mashup, or the text in the second line combination is derived from the first line combination, that is, from the lines of the main film, so the lines in the second line combination can basically match the lines of the main film; each second line combination can include one or more lines, and in the above step S142, the multiple second imitation short sentences corresponding to the multiple second line combinations are determined according to the video clips in which the lines in each second line combination in the multiple second line combinations are located in the main film video. The video may include: when a second line combination includes a line, the video clip corresponding to the line can be determined as the second imitation video corresponding to the second line combination; when a second line combination includes multiple lines, the multiple video clips corresponding to the multiple lines in the second line combination can be spliced ​​according to the starting time and ending time of the video clips corresponding to each line in the second line combination to obtain the second imitation video corresponding to the second line combination; wherein, for each second line combination in the multiple second line combinations, a second imitation short video corresponding to each line combination can be obtained, and multiple second line combinations can obtain multiple second imitation short videos.

[0119] In one possible implementation, a second imitation short video with the top title and / or the fancy words added thereto may be generated with reference to the top title and / or fancy words in the target short video, so as to obtain a second imitation short video with a better promotional effect. Specifically, in the above step S142, the plurality of second imitation short videos corresponding to the plurality of second line combinations are determined based on the video clips in the feature video where each line in each of the plurality of second line combinations is located, including:

[0120] Step S1421: For any second line combination among the plurality of second line combinations, the video segments of each line in the second line combination in the main video are spliced ​​together to obtain an initial imitation short video corresponding to the second line combination;

[0121] Step S1422: Add the top title and / or the fancy word set of the multiple target short videos that match the feature video to the initial imitation short video corresponding to the second line combination, and obtain the second imitation short video corresponding to the second line combination.

[0122] In step S1421, the video segments of the main film containing the lines in the second dialogue combination that match the main film lines can be sequentially spliced ​​based on the start and end times of the video segments in the main film containing the lines in the second dialogue combination, thereby obtaining an initial imitation short video corresponding to the second dialogue combination. If the second dialogue combination includes only one line, the video segment in the main film containing that line can be determined as the initial imitation short video corresponding to the second dialogue combination.

[0123] It should be understood that through the above steps S1214 and S1215, the top title set and / or flower character set of each target short video matching the main video can be obtained. Then, in the above step S1422, since there may be multiple top titles in the top title set of multiple target short videos, a top title can be randomly selected from the multiple top titles for the second line combination and added to the initial imitation short video corresponding to the second line combination. Alternatively, a top title that best matches the line content can be selected from the multiple top titles based on the line content of the second line combination and added to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video with a top title added. This is not limited to the embodiments of the present disclosure. As mentioned above, the flower character set can also record the association relationship between the flower characters identified in the same short video frame and the short video lines. Therefore, based on the association relationship between various flower characters in the flower character set and the short video lines, the flower characters associated with each line in the first line combination can be obtained. Since the text in the second line combination as described above can be understood as a rearranged combination of the original text in the first line combination, the flower characters associated with each line in the second line combination can be determined based on the flower characters associated with each line in the first line combination. Specifically, the flower characters associated with the lines in the second line combination that are the same as those in the first line combination can be determined as the flower characters associated with the lines in the second line combination, so that the flower characters associated with each line in the second line combination can be added to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video with added flower characters.

[0124] Optionally, an artificial intelligence model can also be used to generate top titles and / or fancy words corresponding to each second line combination based on the line content in each second line combination and information such as the plot introduction and plot comments of the main video. The above-mentioned top title set and / or fancy words set can respectively include the top titles and / or fancy words corresponding to each second line combination generated by the artificial intelligence model. Then, in the above-mentioned step S1422, the top titles and / or fancy words corresponding to each second line combination generated by the artificial intelligence model can be added to the initial imitation short video corresponding to each second line combination to obtain a second imitation short video with the added top title and / or fancy words. This is not limited to the embodiments of the present disclosure.

[0125] As described above, background music is usually added to short videos to enhance the emotional expression of the short videos. Therefore, in one possible implementation, step S142, determining the plurality of second imitation short videos corresponding to the plurality of second line combinations based on the video segments in the feature video where the lines in each of the plurality of second line combinations are located, may include:

[0126] Step S1423: For any second line combination among the plurality of second line combinations, if the second line combination includes multiple lines, the video segments containing each line in the second line combination in the main video are spliced ​​together to obtain an initial imitation short video corresponding to the second line combination. The implementation of step S1423 can refer to the above-mentioned step S1421 and is not further described here.

[0127] Step S1424: using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video corresponding to the second line combination based on the line content of the second line combination and the auxiliary information;

[0128] Step S1425: Determine from a preset background music library a second background music that matches the emotion expressed in the initial imitation short video corresponding to the second line combination, and add the second background music to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video corresponding to the second line combination.

[0129] In step S1424, those skilled in the art can use any generative artificial intelligence model known in the art to predict the emotions expressed in the initial imitation short video corresponding to the second line combination based on the line content and auxiliary information of the second line combination. The embodiment of the present disclosure does not limit the type, structure, etc. of the generative artificial intelligence model.

[0130] As mentioned above, plot descriptions and plot reviews of the main film can be obtained from various online platforms such as drama review platforms, social platforms, and video platforms. Plot reviews can include comments with a certain number of likes or reposts, and they can express users' opinions on the highlights of the main film. This can be used to leverage auxiliary information from the main film to help the generative AI model more accurately predict the emotion expressed in the initial imitation short video corresponding to the second line combination based on the content of the lines in the second line combination.

[0131] As described above, each background music in the background music library may carry a business identifier, which is used to describe the usage scenario of the background music, or to describe what short video with what emotion the background music is suitable for. Therefore, in step S1425, the emotion expressed in the initial imitation short video corresponding to the second line combination can be matched with the background music with the business identifier in the background music library to determine the second background music that matches the emotion expressed in the initial imitation short video corresponding to the second line combination from the background music library, and add the second background music to the initial imitation short video corresponding to the second line combination to obtain the second imitation short video of the second line combination.

[0132] As described above, the second line combination may include a line. In this case, a video clip matching the line can be directly determined as the initial imitation short video corresponding to the second line combination, and then the above step S1422 and / or the above steps S1424 to S1425 can be executed on the initial imitation short video to obtain the second imitation short video corresponding to the second line combination.

[0133] It should be understood that for the initial imitation short video corresponding to any second line combination, step S1422 can be simultaneously executed to add a top title and / or flower characters, and steps S1424 to S1425 can be executed to add a matching second background music. That is, the second imitation short video can include the second background music, as well as the top title and / or flower characters; of course, the second imitation short video can also only include the top title and / or flower characters, or only include the second background music, and this embodiment of the present disclosure is not limited to this. For each second line combination in multiple second line combinations, the second imitation short video corresponding to each line combination can be obtained by executing the above steps S1421 to S1422 and / or executing the above steps S1423 to S1425. Multiple second line combinations can obtain multiple second imitation short videos.

[0134] In actual applications, after obtaining multiple first imitation short videos and multiple second imitation short videos, a video list of the generated multiple first imitation short videos and multiple second imitation short videos can be displayed on the front end. The video list can include the video ID, the main video corresponding to the video, the download address, and can also carry the release title of the main video, related topics and related content descriptions and other information, so that users can view the generated multiple first imitation short videos and multiple second imitation short videos, and use the multiple first imitation short videos and multiple second imitation short videos as needed. For example, the multiple first imitation short videos and multiple second imitation short videos can be further used to promote the main video. For example, the multiple first imitation short videos and multiple second imitation short videos can be published on various short video platforms, social platforms and other network platforms, so as to promote the main video to more platform users, which is conducive to increasing the attention and playback volume of the main video and meeting the video platform's promotion and distribution needs for the main video.

[0135] For example, Figure 4 A schematic diagram showing a process of a short video production method according to an embodiment of the present disclosure is shown as follows: Figure 4As shown, the target short video can be obtained and stored in a local database. The video list of the target short videos that have not been downloaded can be obtained, and the target short videos in the video list can be downloaded locally and marked (that is, the target short video can be marked). The short video lines can be identified by OCR technology every time a target short video is downloaded. The short video lines can be in the data format of lines strings and stored in the local database; then the main film lines that match the short video lines or the main film video frames without lines can be matched from the main film lines library to obtain the matching results of the target short video. The matching results can be in the data format of JSON objects, including the starting time and ending time of the main film lines; then, an imitation short video (including the first imitation short video and the second imitation short video) can be generated based on the matching results of the target short video. At the same time, the release title, related topics, content description and other information of the imitation video can be added and stored in the local database, and the target short video can be marked to indicate that the target short video has been produced.

[0136] In actual applications, the short video production method of the above-mentioned embodiment of the present disclosure can use multiple devices to concurrently obtain multiple target short videos and batch produce the first imitation short video and the second imitation short video, which can further improve the efficiency of short video production; it should be understood that data synchronization and task scheduling and other collaborative processes can be performed between multiple devices, and the embodiment of the present disclosure does not limit this.

[0137] According to the method of the embodiment of the present invention, by utilizing a target short video whose user attention level is higher than a specified threshold and matches the main video, a high-quality first imitation short video is generated, which is similar in logic to the target short video, has a strong sense of rhythm, conforms to user preferences, etc. and matches the content of the main video. Then, multiple first imitation short videos that match the main video are used to generate more new second imitation short videos that are similar in logic to the target short video, have a strong sense of rhythm, conform to user preferences, etc. and match the content of the main video. This achieves low-cost and high-efficiency automatic batch production of high-quality imitation short videos. In an exemplary application scenario, it is beneficial to improve the promotional effect of using the first imitation short video and the second imitation short video to promote the main video, and meet the promotional needs of the video platform for the main video.

[0138] In related technologies, although it is possible to select lines that may be highlights by giving all the lines of a feature video in the form of prompts to a generative artificial intelligence model, and then splice the video clips corresponding to these lines to produce imitation short videos; however, the generation effect of this method is difficult to guarantee. Even if the most powerful model on the market is used and enough example data is added to the prompt, it is difficult to guarantee that the selected dialogue lines are in line with the user's preferences.

[0139] The short video production method of the embodiment of the present disclosure identifies the pictures and lines of the target short videos with high user attention on the Internet, as well as important information such as the top title and fancy words in the target short videos, and then matches them with all the lines and pictures in the main video to find the corresponding video clips, and finally splices these video clips in the order in which the lines and pictures appear in the target short video to generate the first imitation short video, so as to achieve the effect of 1:1 imitation of the target short video; after obtaining enough first imitation short videos of target short videos that match a main video, all the lines and pictures of the main video can be matched together to form a first imitation short video. High-quality auxiliary information such as plot comments and plot introductions are given to the generative artificial intelligence model in the form of prompts, so that the generative artificial intelligence model can produce more new line combinations, and then generate multiple new second imitation short videos. This can achieve a 1:N production efficiency and increase the output of short videos. In an exemplary application scenario, a low-cost, batch-produced promotional short video method is realized to meet the promotional demands of various TV dramas, variety shows and other feature video content on the video platform, so that the promotion of imitation short videos can reach more target groups, which is conducive to increasing the playback volume of the feature video.

[0140] Figure 5 A block diagram of a short video production device according to an embodiment of the present disclosure is shown as follows: Figure 5 As shown, the device includes:

[0141] An acquisition module 501 is configured to acquire a plurality of target short videos, wherein the target short videos include short videos whose user attention level is greater than a specified threshold;

[0142] A matching module 502 is configured to match each of the plurality of target short videos with the main video to obtain a matching result for each target short video; wherein the matching result includes: whether the target short video matches the main video, and, when the target short video matches the main video, segment information of a video segment in the main video that matches the target short video;

[0143] The first imitation module 503 is configured to determine, for any target short video among the multiple target short videos, a first imitation short video of the target short video based on segment information of a video segment in the main video that matches the target short video, if the target short video matches the main video;

[0144] The second imitation module 504 is used to generate multiple second imitation short videos based on multiple first imitation short videos of multiple target short videos matching the main video; the multiple first imitation short videos and the multiple second imitation short videos are used to promote the main video.

[0145] In a possible implementation, the clip information includes the main film lines in the video clip, as well as the start time and the end time of the video clip; the matching of each target short video in the multiple target short videos with the main film video to obtain the matching result of each target short video includes: for any target short video in the multiple target short videos, extracting short video lines from the target short video to obtain a first line set corresponding to the target short video, the first line set including multiple short video lines appearing in the target short video; matching each short video line in the first line set with the main film lines in the main film line library; The character string is matched to obtain a line matching result of each short video line; wherein the main film line library includes all the main film lines in the main film video and the start time and end time of the video clip containing each main film line in all the main film lines; the line matching result includes: whether it matches the main film lines in the main film line library, and, when matching the main film lines in the main film line library, the start time and end time of the video clip containing the matched main film lines; the matching result of the target short video is determined based on the line matching result of each short video line in the first line set.

[0146] In one possible implementation, the segment information includes a main video frame in the video segment, and the start and end times of the video segment in which the main video frame is located. Matching each target short video in the multiple target short videos with the main video to obtain a matching result for each target short video includes: for any target short video in the multiple target short videos, if no short video lines are extracted from the target short video, or if the number of short video lines extracted from the target short video is less than a specified number of lines, performing similarity matching between the short video frames in the target short video and the main video frames in the main frame library to obtain a frame matching result for the short video frames in the target short video; wherein the main frame library includes all main video frames extracted at intervals from the main video; the frame matching result includes: whether the frame matches the main video frame in the main frame library, and, if the frame matches the main video frame in the main frame library, the frame time of the matched main video frame; and determining the matching result of the target short video based on the frame matching result of the short video frames in the target short video.

[0147] In a possible implementation, the feature video includes multiple video episodes, and the method of performing string matching on each short video line in the first line set with the feature film lines in the feature film line library to obtain a line matching result for each short video line includes: for any short video line in the first line set, performing string matching on the short video lines with the feature film lines in the feature film line library to obtain multiple feature film lines that match the short video lines, and then selecting, based on a target set video where the context lines of the short video lines are located, feature film lines that appear in the target set video from the multiple feature film lines that match the short video lines; and determining the line matching result for the short video lines based on the start time and end time of the video clip where the selected feature film lines that appear in the target set video are located.

[0148] In a possible implementation, the performing string matching of each short video line in the first line set with the main film lines in the main film line library to obtain the line matching results of each short video line includes: for any short video line in the first line set, performing string matching of the short video lines with the main film lines in the main film line library to obtain multiple main film lines that match the short video lines, performing similarity matching of the main film video frames containing the multiple main film lines that match the short video lines with the short video frames containing the short video lines to obtain the main film video frames with the highest similarity to the short video frames containing the short video lines; and determining the line matching results of the short video lines based on the main film lines corresponding to the main film video frame with the highest similarity to the short video frame containing the short video lines and the start time and end time of the video clip containing the short video lines.

[0149] In one possible implementation, the extracting of short video lines from a target short video to obtain a first line set corresponding to the target short video includes: extracting short video frames from the target short video at n frames per second to obtain multiple extracted short video frames, and performing character recognition on each of the short video frames extracted from the target short video to obtain recognized character strings and positions of the character strings in the multiple short video frames, where n is a positive integer; clustering and grouping the recognized character strings in the multiple short video frames according to the positions of the recognized character strings in the multiple short video frames to obtain multiple character string groups; and determining a first character string group containing the largest number of character strings among the multiple character string groups as the first line set.

[0150] In a possible implementation, the first imitation short video of the target short video is determined based on the segment information of the video segment that matches the target short video in the main video, including: when there are multiple video segments that match the target short video in the main video, according to the segment information of the multiple video segments that match the target short video in the main video, the multiple video segments that match the target short video in the main video are spliced ​​to obtain an initial imitation short video corresponding to the target short video; according to the top title set and / or flower character set of the target short video, the top title and / or flower character set are added to the initial imitation short video corresponding to the target short video to obtain the first imitation short video of the target short video.

[0151] In one possible implementation, the device further includes: a title determination module, configured to determine a second character string group, among the multiple character string groups, whose character strings are located at the top of the short video frame and whose character strings are repeated more than a specified threshold number of times, as a top title set of the target short video; and / or, a flower character determination module, configured to determine a character string group, among the multiple character string groups, other than the first character string group and the second character string group, as a flower character set in the target short video.

[0152] In a possible implementation, the first imitation short video of the target short video is determined based on the segment information of the video segment that matches the target short video in the main video, including: when there are multiple video segments that match the target short video in the main video, the multiple video segments that match the target short video in the main video are spliced ​​according to the segment information of the multiple video segments that match the target short video in the main video to obtain an initial imitation short video corresponding to the target short video; using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video according to the lines corresponding to the initial imitation short video and the auxiliary information of the main video; the auxiliary information includes a plot introduction and / or a plot commentary; determining the first background music that matches the emotion expressed in the initial imitation short video from a preset background music library, and adding the first background music to the initial imitation short video to obtain the first imitation short video of the target short video.

[0153] In one possible implementation, the generating of multiple second imitation short videos based on multiple first imitation short videos of multiple target short videos matching the main video includes: when more than a specified number of multiple first imitation short videos of multiple target short videos matching the main video are obtained, generating multiple second line combinations based on the first line combinations in the multiple first imitation short videos and auxiliary information related to the main video using a generative artificial intelligence model; the auxiliary information includes plot introduction and / or plot commentary; determining the video segment in which each line in each second line combination in the multiple second line combinations is located in the main video, and determining the multiple second imitation short videos corresponding to the multiple second line combinations based on the video segment in which each line in each second line combination in the multiple second line combinations is located in the main video.

[0154] In a possible implementation, determining the video segment in the feature video where each line in each of the multiple second line combinations is located includes: for any second line combination in the multiple second line combinations, performing string matching on each line in the second line combination with the feature lines in the feature line library to obtain a line matching result for each line in the second line combination, the line matching result including: whether it matches the feature lines in the feature line library, and, when matching the feature lines in the feature line library, the start time and end time of the matched feature lines and the video segment where the matched feature lines are located; and determining the video segment in the feature video where each line in the second line combination is located based on the line matching result of each line in the second line combination.

[0155] In a possible implementation, the method of determining a plurality of second imitation short videos corresponding to the plurality of second line combinations based on the video segments in the feature video where each line in each second line combination in the plurality of second line combinations is located includes: for any second line combination in the plurality of second line combinations, when the second line combination includes multiple lines, splicing the video segments in the feature video where each line in the second line combination is located to obtain an initial imitation short video corresponding to the second line combination; and adding a top title and / or a flower character set to the initial imitation short video corresponding to the second line combination based on the top title set and / or flower character set of a plurality of target short videos matching the feature video to obtain a second imitation short video corresponding to the second line combination.

[0156] In a possible implementation, the method of determining a plurality of second imitation short videos corresponding to the plurality of second line combinations based on the video segments in the feature video where each line in each second line combination in the plurality of second line combinations is located includes: for any second line combination among the plurality of second line combinations, when the second line combination includes multiple lines, splicing the video segments in the feature video where each line in the second line combination is located to obtain an initial imitation short video corresponding to the second line combination; using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video corresponding to the second line combination based on the line content of the second line combination and the auxiliary information; determining second background music from a preset background music library that matches the emotion expressed in the initial imitation short video corresponding to the second line combination, and adding the second background music to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video corresponding to the second line combination.

[0157] According to the device of the embodiment of the present disclosure, by utilizing the target short video whose user attention degree is higher than the specified threshold and matches the main video, a high-quality first imitation short video is generated, which is similar in logic to the target short video, has a strong sense of rhythm, conforms to user preferences, etc. and matches the content of the main video. Then, multiple first imitation short videos that match the main video are used to generate more new second imitation short videos that are similar in logic to the target short video, have a strong sense of rhythm, conform to user preferences, etc. and match the content of the main video, thereby achieving low-cost and high-efficiency automatic batch production of high-quality imitation short videos. In an exemplary application scenario, it is beneficial to improve the promotional effect of using the first imitation short video and the second imitation short video to promote the main video, and meet the promotional needs of the video platform for the main video.

[0158] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0159] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0160] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0161] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0162] Figure 6 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 6 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0163] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0164] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0165] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0166] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0167] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0168] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, device instructions, device-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, which may be executed by the computer-readable program instructions to implement various aspects of the present disclosure.

[0169] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0170] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a device, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0171] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0172] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0173] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A short video production method, characterized in that: include: Acquire multiple target short videos, where the target short videos include short videos whose user attention level is greater than a specified threshold; Matching each target short video among the multiple target short videos with the main video to obtain a matching result for each target short video; wherein the matching result includes: whether the target short video matches the main video, and, when the target short video matches the main video, segment information of the video segment in the main video that matches the target short video; For any target short video among the multiple target short videos, if the target short video matches the main video, determining a first imitation short video of the target short video based on segment information of a video segment in the main video that matches the target short video, wherein the segment information includes the main film lines in the video segment, and the start time and end time of the video segment; A plurality of second imitation short videos are generated based on a plurality of first imitation short videos of a plurality of target short videos that match the feature video.

2. The method according to claim 1, characterized in that Matching each target short video in the plurality of target short videos with the feature video to obtain a matching result for each target short video includes: For any target short video among the multiple target short videos, extract short video lines from the target short video to obtain a first line set corresponding to the target short video, where the first line set includes multiple short video lines appearing in the target short video; Performing string matching on each short video line in the first line set with the main film lines in the main film line library to obtain a line matching result for each short video line; wherein the main film line library includes all main film lines in the main film video and the start time and end time of the video clip containing each main film line in all the main film lines; the line matching result includes: whether the main film lines in the main film line library are matched, and, if the main film lines in the main film line library are matched, the start time and end time of the video clip containing the matched main film lines; The matching result of the target short video is determined according to the line matching results of each short video line in the first line set.

3. The method according to claim 2, characterized in that The segment information includes the main video frame in the video segment, and the start time and end time of the video segment where the main video frame is located. Matching each target short video in the plurality of target short videos with the feature video to obtain a matching result for each target short video includes: For any target short video among the multiple target short videos, when no short video lines are extracted from the target short video, or when the number of short video lines extracted from the target short video is less than a specified number of lines, similarity matching is performed on the short video frames in the target short video and the main video frames in the main film frame library to obtain a frame matching result of the short video frames in the target short video; wherein the main film frame library includes all main film video frames extracted at intervals from the main film video; the frame matching result includes: whether the frame matches the main film video frames in the main film frame library, and, when the frame matches the main film video frames in the main film frame library, the frame time of the matched main film video frames; According to the frame matching results of the short video frames in the target short video, the matching result of the target short video is determined.

4. The method according to claim 2, characterized in that The feature video includes multiple episodes of video, and performing string matching on each short video line in the first line set with the feature lines in the feature line library to obtain a line matching result for each short video line includes: For any short video line in the first line set, after performing string matching on the short video line with the main film lines in the main film line library to obtain multiple main film lines that match the short video line, select the main film lines that appear in the target set video from the multiple main film lines that match the short video line based on the target set video where the context lines of the short video line are located; The lines matching result of the short video lines is determined according to the start time and the end time of the video segment containing the selected main film lines appearing in the target set of videos.

5. The method according to claim 2, characterized in that The performing string matching of each short video line in the first line set with the main film lines in the main film line library to obtain a line matching result of each short video line includes: For any short video line in the first line set, after performing string matching on the short video line and the main film lines in the main film line library to obtain multiple main film lines matching the short video line, performing similarity matching on the main film video frames containing the multiple main film lines matching the short video line and the short video frame containing the short video line to obtain the main film video frame having the highest similarity to the short video frame containing the short video line; The lines matching result of the short video lines is determined based on the main film lines corresponding to the main film video frame with the highest similarity to the short video frame where the short video lines are located and the start time and end time of the video clip where the short video lines are located.

6. The method according to claim 2, characterized in that The step of extracting short video lines from a target short video to obtain a first line set corresponding to the target short video includes: Extract short video frames from the target short video at n frames per second to obtain multiple extracted short video frames, and perform character recognition on each of the short video frames extracted from the target short video to obtain recognized character strings and positions of the character strings in the multiple short video frames, where n is a positive integer; Clustering and grouping the character strings identified in the multiple short video frames according to positions of the character strings identified in the multiple short video frames to obtain multiple character string groups; A first character string group including the largest number of character strings among the plurality of character string groups is determined as the first line set.

7. The method according to claim 6, characterized in that The step of determining a first imitation short video of the target short video based on the segment information of the video segment matching the target short video in the feature video includes: In the case where there are multiple video segments in the main video that match the target short video, the multiple video segments in the main video that match the target short video are spliced ​​together according to the segment information of the multiple video segments in the main video that match the target short video to obtain an initial imitation short video corresponding to the target short video; According to the top title set and / or fancy word set of the target short video, the top title and / or fancy word set are added to the initial imitation short video corresponding to the target short video to obtain a first imitation short video of the target short video.

8. The method according to claim 7, characterized in that The method further comprises: Determine a second character string group in which the character string is located at the top of the short video frame and the number of repetitions of the character string exceeds a specified threshold as the top title set of the target short video; and / or, The character string groups other than the first character string group and the second character string group in the multiple character string groups are determined as the flower character set in the target short video.

9. The method according to any one of claims 1 to 8, characterized in that The step of determining a first imitation short video of the target short video based on the segment information of the video segment matching the target short video in the feature video includes: In the case where there are multiple video segments in the main video that match the target short video, the multiple video segments in the main video that match the target short video are spliced ​​together according to the segment information of the multiple video segments in the main video that match the target short video to obtain an initial imitation short video corresponding to the target short video; Using a generative artificial intelligence model to predict the emotion expressed in the initial imitation short video based on the lines corresponding to the initial imitation short video and auxiliary information of the feature video; the auxiliary information includes a plot introduction and / or plot commentary; A first background music that matches the emotion expressed in the initial imitation short video is determined from a preset background music library, and the first background music is added to the initial imitation short video to obtain a first imitation short video of the target short video.

10. The method according to claim 1, characterized in that The step of generating a plurality of second imitation short videos based on a plurality of first imitation short videos of a plurality of target short videos that match the main video includes: When a plurality of first imitation short videos are obtained that are more than a specified number of target short videos that match the main video, a plurality of second dialogue combinations are generated using a generative artificial intelligence model based on the first dialogue combinations in the plurality of first imitation short videos and auxiliary information related to the main video; the auxiliary information includes plot introductions and / or plot comments; Determine the video segment in which each line in each of the multiple second line combinations is located in the feature video, and determine the multiple second imitation short videos corresponding to the multiple second line combinations based on the video segment in which each line in each of the multiple second line combinations is located in the feature video.

11. The method according to claim 10, characterized in that The step of determining the video segment in which each line in each of the plurality of second line combinations is located in the feature video includes: For any second line combination among the plurality of second line combinations, performing string matching on each line in the second line combination with lines from the main film in the main film line library to obtain a line matching result for each line in the second line combination, the line matching result including: whether the line matches the main film line in the main film line library, and, if the line matches the main film line in the main film line library, the start time and end time of the video clip containing the matched main film line; According to the line matching results of each line in the second line combination, the video segment where each line in the second line combination is located in the feature video is determined.

12. The method according to claim 10, characterized in that The step of determining, based on the video segments in the feature video where each line in each of the plurality of second line combinations is located, a plurality of second imitation short videos corresponding to the plurality of second line combinations includes: For any second line combination among the plurality of second line combinations, if the second line combination includes multiple lines, splicing the video segments containing each line in the second line combination in the feature video to obtain an initial imitation short video corresponding to the second line combination; According to the top title set and / or flower character set of multiple target short videos matching the feature video, the top title and / or flower character are added to the initial imitation short video corresponding to the second line combination to obtain the second imitation short video corresponding to the second line combination.

13. The method according to any one of claims 10 to 12, characterized in that: The step of determining, based on the video segments in the feature video where each line in each of the plurality of second line combinations is located, a plurality of second imitation short videos corresponding to the plurality of second line combinations includes: For any second line combination among the plurality of second line combinations, if the second line combination includes multiple lines, splicing the video segments containing each line in the second line combination in the feature video to obtain an initial imitation short video corresponding to the second line combination; Using a generative artificial intelligence model, based on the content of the second line combination and the auxiliary information, predict the emotion expressed in the initial imitation short video corresponding to the second line combination; Determine from a preset background music library the second background music that matches the emotion expressed in the initial imitation short video corresponding to the second line combination, and add the second background music to the initial imitation short video corresponding to the second line combination to obtain a second imitation short video corresponding to the second line combination.

14. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 13 when executing the instructions stored in the memory.

15. A computer program product, characterized in that The invention comprises a computer-readable code, or a non-volatile computer-readable storage medium carrying a computer-readable code, wherein when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Video association method and device, server and readable storage medium

    CN111767796A

  • Video generation method and device, computer equipment and storage medium

    CN117278699A

  • Explanation information generation method, device, equipment, medium and program product

    CN117793478A