A method and system for quickly batch generating short videos with a high completion rate
By intelligently extracting Internet hot content and optimizing editing methods with genetic algorithms and prediction models, the problem of lack of personalization and targeted content in the existing short video batch generation methods is solved, and the completion rate and commercial conversion capabilities are significantly improved.
Patent Information
- Application Number
- CN202510145638.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing short video batch generation methods are difficult to adapt to the target market demand, and the generated content lacks personalization and targeting, resulting in low completion rate and unsatisfactory traffic monetization effect.
By intelligently extracting hot content on the Internet, combining genetic algorithms and prediction models to optimize the editing method of short videos, we generate personalized short videos that are highly consistent with user interests and hot trends.
It significantly improves the completion rate of short videos, enhances its commercial conversion capabilities, adapts to the fierce market traffic competition needs, and provides support for enterprises to achieve efficient traffic monetization.
Smart Images

Figure CN119629441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short videos, and more specifically, it relates to a method and system for quickly batch generating short videos with a high completion rate. Background Art
[0002] With the rapid development of the new media industry, short videos have become an important tool for traffic monetization due to their low production cost, fast dissemination speed, diverse content forms, etc. Enterprises increasingly rely on the play volume and completion rate of short videos as the core means of traffic monetization.
[0003] Existing methods for batch generating short videos mainly adopt templatization and automation technologies, simply splicing video, image, text, audio and other materials to quickly produce a large number of short videos. However, this fixed-template method is difficult to meet the needs of the target market, and the generated content lacks personalization and pertinence, unable to effectively attract target users, resulting in a low completion rate and unsatisfactory traffic monetization effect. Especially in the fierce market traffic competition, this method is difficult to meet the traffic monetization needs of new media enterprises. Summary of the Invention
[0004] The present invention provides a method and system for quickly batch generating short videos with a high completion rate, which solves the technical problems proposed in the background art.
[0005] The present invention provides a method for quickly batch generating short videos with a high completion rate, including:
[0006] Step 1, within a first preset time period, obtain the first hot keywords of the Internet platform at fixed time intervals, and obtain the first hot videos corresponding to the first hot keywords;
[0007] Step 2, construct a video material library, which includes a text unit, a background music unit and a dubbing unit; the text unit is used to generate text for the first hot video, the background music unit is used to generate background music for the first hot video, and the dubbing unit is used to generate dubbing according to the generated text;
[0008] Step 3, based on the first hot video, sequentially construct a random population and an optimization population, and obtain a target individual according to the optimization population;
[0009] Step 4, loop step 3 to obtain a number of target individuals and corresponding objective function values, combine the target individuals and the corresponding first hot keywords as training samples, and use the corresponding objective function values as sample labels to train an individual prediction model;
[0010] Step 5, within a second preset time period, obtain the second hot keywords, and based on the genetic algorithm and the individual prediction model, obtain the editing method of the second hot video to generate the corresponding short video.
[0011] Further, the text unit, including a video parsing tool and an insertion tool, includes:
[0012] A video parsing tool for inputting a first hot video and extracting the dialogue text of the first hot video;
[0013] An insertion tool for inserting connection text between the dialogue texts, and the connection text includes:
[0014] Step 21: Obtain the hot text corresponding to the first hot keyword, and disassemble the hot text based on a language model to obtain the event cause text, process text, and result text of the hot text;
[0015] Step 22: Obtain the timestamps corresponding to the dialogue text of the first hot video. If the time interval between adjacent dialogue texts is greater than a preset time, the interval between adjacent dialogue texts is used as an insertion candidate interval;
[0016] Step 23: Map the cause text, process text, and result text, and the dialogue texts at both ends of the insertion candidate interval into word vectors respectively, and calculate the correlation degree between the cause text, process text, and result text and the insertion candidate interval based on the word vectors. The calculation formula is as follows:
[0017] ;
[0018] Wherein, represents the correlation degree between the cause text, process text, or result text and the insertion candidate interval, represents the word vector of the cause text, process text, or result text, represents the word vector of the dialogue text on the front side of the insertion candidate interval, represents the word vector of the dialogue text on the back side of the insertion candidate interval, , and respectively represent the first weight parameter, the second weight parameter, and the third weight parameter, represents calculating the Euclidean norm, represents calculating the inner product, represents calculating the square of the Euclidean norm;
[0019] Step 24: Respectively determine the insertion candidate intervals with the maximum correlation degree for the cause text, process text, and result text, and allocate the cause text, process text, and result text as connection texts to the insertion candidate intervals with the maximum correlation degree one by one.
[0020] Further, the dubbing unit and the background music unit are as follows:
[0021] Obtain open-source background music materials and dubbing voices, and cluster the background music materials and dubbing voices based on a preset emotional style to obtain several emotional background music materials and several emotional dubbing voices;
[0022] Perform emotional recognition on the first hot keyword to obtain the emotional characteristics of the first hot keyword, and match the corresponding emotional background music materials and emotional dubbing voices based on the emotional characteristics.
[0023] Further, construct a random population, including:
[0024] Represent both the open-source background music materials and dubbing voices through real number coding;
[0025] Perform writing style conversion on the connecting text through a language model, and represent the writing style through real number coding;
[0026] According to the first hot video, initialize several random individuals that meet the constraint conditions. The coding of the random individuals is as follows:
[0027] ; where represents the i-th writing style, represents the i-th background music material, represents the i-th dubbing voice;
[0028] The constraint conditions include: Belong to the emotional background music materials corresponding to the first hot video, Belong to the emotional dubbing voices corresponding to the first hot video;
[0029] Construct an objective function, which is used to evaluate the short video obtained by randomly editing the first hot video. The formula of the objective function is as follows:
[0030] ;
[0031] ;
[0032] ;
[0033] where represents the objective function value, represents the time weight, represents the lifespan of the short video, 、 and respectively represent the vectors of the lifespan at the 、 and th time points, represents the ratio of the number of likes to the number of views at the h-th time point, represents the ratio of the number of complete plays to the number of plays at the h-th time point, represents the cross product of vectors, , and represent the curvature weight, weighted play weight, and interaction weight respectively, , and are all non-zero, and , and sum to 1, represents the preset end time point of the lifeline, represents the start time point of the lifeline, represents the h-th time point of the lifeline, the g-th time point of the lifeline, represents the number of time points of the lifeline, represents the first index of the number of time points of the lifeline, represents the second index of the number of time points of the lifeline, is not equal to , represents the vector of the lifeline at the n-th time point;
[0034] Set the target threshold, calculate the objective function value for each random individual. If the objective function value of the random individual > the target threshold, then take the corresponding random individual as a high-quality individual.
[0035] Furthermore, construct an optimization population, including:
[0036] Initialize and generate the optimization population based on the encoding of several high-quality individuals, specifically as follows:
[0037] Combine the writing style encoding, background music material encoding, and voice-over tone encoding of high-quality individuals respectively to obtain a writing style encoding library, a background music material encoding library, and a voice-over tone encoding library;
[0038] The writing style encoding of the optimization individuals in the optimization population is randomly selected from the writing style encoding library, the background music material encoding is randomly selected from the background music material encoding library, and the voice-over tone encoding is randomly selected from the voice-over tone encoding library to obtain several optimization individuals;
[0039] Calculate the objective function values of several optimization individuals based on the objective function, and take the optimization individual with the largest objective function value as the target individual.
[0040] Furthermore, train to obtain an individual prediction model, including:
[0041] Map the first hot keyword obtained within the first preset time period to a hot word vector, and merge the hot word vector and the corresponding target individual encoding to obtain a feature vector;
[0042] Input the feature vector into the individual prediction model, and the output of the individual prediction model represents the predicted evaluation score of the short video obtained by editing the first hot video corresponding to the first hot keyword according to the target individual encoding;
[0043] Take the mean squared error between the predicted evaluation score and the target function value corresponding to the target individual as the loss function of the individual prediction model, and update the hyperparameters of the hidden layer of the individual prediction model through backpropagation.
[0044] Further, the hidden layer of the individual prediction model is specifically as follows:
[0045] ;
[0046] ;
[0047] Among them, represents the mapping output of the feature vector in the th hidden layer, represents the mapping output of the feature vector in the th hidden layer, represents the activation function, represents element-wise multiplication, represents based on constructed left diagonal matrix, and the left diagonal matrix represents a sparse matrix with as the left diagonal element of the matrix and the rest of the elements being 0, represents based on constructed right diagonal matrix, and the right diagonal matrix represents a sparse matrix with as the right diagonal element of the matrix and the rest of the elements being 0, represents the th hidden layer feature space, represents the bias parameter of the th hidden layer feature space, represents the transpose operation.
[0048] Further, based on the genetic algorithm and the individual prediction model, obtain the editing method of the second hot video, including:
[0049] Step 81, map the second hot keyword obtained in the second preset time period to a hot word vector ;
[0050] Step 82, initialize a number of genetic individuals, and the genetic individuals are represented as: ; wherein, , and respectively represent the th writing style, the th music score material, and the th dubbing timbre;
[0051] Step 83: Use the individual prediction model as the genetic adaptation function of the genetic population, and obtain the genetic adaptation function value of each genetic individual through the individual prediction model;
[0052] Step 84: Based on the genetic adaptation function values, sort each genetic individual from largest to smallest to obtain a feature ranking;
[0053] Step 85: Retain a preset number of genetic individuals in the feature ranking from front to back, and use the remaining individuals as the parent generation for update. The update includes:
[0054] Extract a random number X for each parent generation within a fixed value range. If X is greater than the first preset threshold P1, randomly select a parent generation for one-way crossover for the parent generation;
[0055] If X is less than the second preset threshold P2, mutate the parent generation;
[0056] If P2 < X < P1, suspend the update of the parent generation;
[0057] Step 86: If the loop of Step 84 and Step 85 reaches the preset number of times, stop the update, output the genetic individual with the largest genetic adaptation function value, obtain the editing method of the second hot video corresponding to the second hot keyword according to this genetic individual, and generate the corresponding short video.
[0058] A system for quickly batch producing short videos with a high completion rate, applied to the described method, includes:
[0059] A data acquisition module, configured to obtain the first hot keyword on the Internet platform at fixed time intervals within the first preset time period, and obtain the first hot video corresponding to the first hot keyword;
[0060] A clip library module, configured to build a video material library. The video material library includes a text unit, a music score unit, and a dubbing unit; the text unit is used to generate text for the first hot video, the music score unit is used to generate music score for the first hot video, and the dubbing unit is used to generate dubbing according to the generated text;
[0061] A population optimization module, configured to sequentially build a random population and an optimization population based on the first hot video, and obtain a target individual according to the optimization population;
[0062] An individual prediction module, which is used to obtain a number of target individuals and corresponding objective function values, combine the target individuals and corresponding first hot keywords as training samples, and use the corresponding objective function values as sample labels to train an individual prediction model;
[0063] A clip planning module, which, within a second preset time period, obtains second hot keywords, and based on a genetic algorithm and the individual prediction model, obtains a clip method for a second hot video to generate a corresponding short video.
[0064] The beneficial effects of the present invention are as follows: By intelligently extracting Internet hot content and combining a genetic algorithm and a prediction model to optimize the clip method of short videos, personalized short videos that highly fit user interests and hot trends can be quickly generated. This method not only improves the completion rate of short videos but also significantly enhances their commercial conversion ability, adapts to the current market demand with fierce traffic competition, and provides strong support for enterprises to achieve efficient traffic monetization. Brief Description of the Drawings
[0065] Figure 1 It is a flowchart of a method for quickly batch generating short videos with a high completion rate according to the present invention. Detailed Embodiments
[0066] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.
[0067] As Figure 1 shown, a method for quickly batch generating short videos with a high completion rate includes:
[0068] Step 1, within a first preset time period, obtain first hot keywords of an Internet platform at fixed time intervals, and obtain first hot videos corresponding to the first hot keywords;
[0069] Step 2, construct a video material library, which includes a text unit, a background music unit, and a voiceover unit; the text unit is used to generate text for the first hot video, the background music unit is used to generate background music for the first hot video, and the voiceover unit is used to generate a voiceover according to the generated text;
[0070] Step 3, based on the first hot video, sequentially construct a random population and an optimization population, and obtain target individuals according to the optimization population;
[0071] Step 4, repeat Step 3 to obtain a number of target individuals and corresponding objective function values. Combine the target individuals and the corresponding first hot keywords as training samples, and use the corresponding objective function values as sample labels to train an individual prediction model.
[0072] Step 5, within a second preset time period, obtain second hot keywords, and based on the genetic algorithm and the individual prediction model, obtain the editing method of the second hot video to generate a corresponding short video.
[0073] In an embodiment of the present invention, the text unit includes a video parsing tool and an insertion tool, and includes:
[0074] The video parsing tool is used to input the first hot video and extract the dialogue text of the first hot video.
[0075] The insertion tool is used to insert connection text between the dialogue texts. The connection text includes:
[0076] Step 21, obtain the hot text corresponding to the first hot keyword, and based on the language model, disassemble the hot text to obtain the event cause text, process text, and result text of the hot text.
[0077] Step 22, obtain the time stamps corresponding to the dialogue texts of the first hot video. If the time interval between adjacent dialogue texts is greater than the preset time, the interval between adjacent dialogue texts is used as the insertion candidate interval.
[0078] Step 23, map the cause text, process text, and result text, and the dialogue texts at both ends of the insertion candidate interval into word vectors respectively, and calculate the correlation degree between the cause text, process text, and result text and the insertion candidate interval based on the word vectors. The calculation formula is as follows:
[0079] ;
[0080] Wherein, represents the correlation degree between the cause text, process text, or result text and the insertion candidate interval, represents the word vector of the cause text, process text, or result text, represents the word vector of the dialogue text on the front side of the insertion candidate interval, represents the word vector of the dialogue text on the rear side of the insertion candidate interval, 、 and respectively represent the first weight parameter, the second weight parameter, and the third weight parameter, represents calculating the Euclidean norm, represents calculating the inner product, represents calculating the square of the Euclidean norm;
[0081] Step 24: Determine the insertion candidate intervals with the highest correlation degrees for the cause text, process text, and result text respectively, and allocate the cause text, process text, and result text as connection texts to the insertion candidate intervals with the highest correlation degrees one by one.
[0082] It should be noted that the dubbed voice timbre is configured for the connection texts, that is, all connection texts are played through the dubbed voice timbre.
[0083] Specifically, in the short video dissemination, the first hot video often lacks sufficient information background. It is difficult for the audience to fully understand the cause, process, and result of the event only relying on the pictures and scattered dialogues. This may lead to insufficient understanding of the video content by the audience, and then affect the viewing experience and the completion rate of the video. To solve this problem, in this embodiment, by generating and inserting connection texts, the video content logic is supplemented and improved, and a clearer event context is provided for the audience. The dialogue text is the core part for the audience to directly obtain information from the video, and its extraction and timestamp marking provide a basis for the subsequent planning of inserting texts. Obtain the hot texts related to the first hot keyword (such as news reports, event introductions, etc.). Based on the language model, these texts are structurally disassembled into three parts: "event cause", "process", and "result". Each part of the text corresponds to the background of the event occurrence, the main process, and the final result, ensuring clear logic and complete information. According to the dialogue text timestamps obtained by the video parsing tool, analyze the time intervals between adjacent dialogues. If the interval is greater than the set threshold (that is, there is an obvious breakpoint between the dialogues), mark this interval as a potential insertion point. Convert the generated cause text, process text, result text, and the dialogue texts at both ends of the interval to be inserted into word vectors respectively. Use a mathematical model to calculate the semantic correlation degree between the connection text and the context dialogue text, ensuring that the inserted connection text is semantically consistent with the front and back. Calculate the correlation degrees for the cause, process, and result texts respectively, and select the interval with the highest correlation degree as the insertion point. According to the time logic of the video, insert the three types of connection texts (cause, process, result) into the best positions respectively to make the video logic clear and the transition natural.
[0084] It should be noted that the video parsing tool can be Whisper, and the language model can be Chatgpt.
[0085] For example, if there are 10 dialogue texts, then 9 intervals of dialogue texts are obtained. It is judged that 4 intervals of dialogue texts are greater than 10 seconds, then 4 insertion candidate intervals are obtained. Insert the cause text, process text, and result text into the 4 intervals of dialogue texts. Based on the calculation of the correlation degree, it is obtained that the cause text matches the first insertion candidate interval, the process text matches the third insertion candidate interval, and the result text matches the fourth insertion candidate interval. Then, insert the cause text, process text, and result text according to the matching results.
[0086] In one embodiment of the present invention, the dubbing unit and the background music unit are as follows:
[0087] Obtain open-source background music materials and dubbing voices, and cluster the background music materials and dubbing voices based on a preset emotional style to obtain several emotional background music materials and several emotional dubbing voices;
[0088] Perform emotional recognition on the first hot keyword to obtain the emotional characteristics of the first hot keyword, and match the corresponding emotional background music materials and emotional dubbing voices based on the emotional characteristics.
[0089] Specifically, the dubbing and background music of short videos are important factors affecting the user viewing experience. Good sound effects can not only enhance the emotional appeal of the video but also effectively improve the audience's attention and engagement. Through the intelligent emotional classification of background music materials and dubbing timbres, as well as the precise matching with the emotional characteristics of the first hot keyword, the present invention achieves a high emotional fit between the content of short videos and the sound effects, enhancing the attractiveness and completion rate of the videos. By collecting open-source background music material libraries (such as audio clips provided by open-source platforms) and dubbing timbres (such as voice samples with different voices and intonations in the timbre library), rich sound effect resources are provided. These materials are classified and standardized to ensure the efficient use of subsequent models. Emotional classification criteria: For background music emotion classification, categories such as happy, sad, exciting, calm, etc. For dubbing timbre classification, categories such as lively, solemn, humorous, sentimental, etc. Use an emotion recognition algorithm (which can be an Autoencoder based on deep learning) to extract and cluster the characteristics of background music materials and dubbing timbres. Feature extraction can be based on audio signal processing (such as pitch, rhythm, timbre) and emotional labels (such as manual or algorithmic emotional annotation of audio). Obtain several groups of emotional background music materials and dubbing timbre groups, which can be directly used for emotional matching with video content. Conduct emotion recognition on the first hot keyword, inputting the keyword extracted from short videos or hot events. An emotion analysis tool, an emotion analysis model based on NLP (Natural Language Processing) (Chatgpt language model), classifies the keyword. The classification result can be multiple emotion labels (such as positive emotion, negative emotion, neutral emotion), or more fine-grained emotion types (such as anger, joy, sadness, etc.). Emotion feature extraction: Use the emotion characteristics of the keyword (such as positive or negative tendency or specific emotion type) as the matching basis to generate the most suitable sound effects for the video content. According to the emotion characteristics of the keyword, select materials with consistent emotions from the pre-clustered background music materials and dubbing timbres. For example, if the emotion characteristic of the keyword is "joy", then match "lively background music" and "lively dubbing timbre"; if the emotion characteristic is "sadness", then match "soft background music" and "low-pitched dubbing timbre". Add the matched background music materials and dubbing timbres to the short video to make the emotional atmosphere of the overall audio-visual content more consistent.
[0090] In an embodiment of the present invention, constructing a random population includes:
[0091] Both the open-source background music materials and dubbing timbres are represented by real number encoding;
[0092] The transition text is subjected to writing style conversion through a language model, and the writing style is represented by real number encoding;
[0093] According to the first hot video, initialize a number of random individuals that meet the constraint conditions. The random individual encoding is as follows:
[0094] ; where, represents the i-th writing style, represents the i-th background music material, represents the i-th dubbing timbre;
[0095] The constraint conditions include: belongs to the emotional background music material corresponding to the first hot video, belongs to the emotional dubbing timbre corresponding to the first hot video;
[0096] Construct an objective function, which is used to evaluate the short video obtained by randomly individual editing the first hot video. The formula of the objective function is as follows:
[0097] ;
[0098] ;
[0099] ;
[0100] where, represents the objective function value, represents the time weight, represents the lifeline of the short video, , and respectively represent the vectors of the lifeline at the , and th time points, represents the ratio of the number of likes to the number of views at the h-th time point, represents the ratio of the number of complete plays to the number of views at the h-th time point, represents the cross product of vectors, , and respectively represent the curvature weight, the weighted play weight and the interaction weight, , and are all not zero, and , and sum to 1, represents the preset end time point of the lifeline, represents the start time point of the lifeline, represents the h-th time point of the lifeline, the g-th time point of the lifeline, represents the number of time points of the lifeline, represents the first index of the number of time points of the lifeline, represents the second index of the number of time points of the lifeline, not equal to , represents the vector of the lifeline at the nth time point;
[0101] Set a target threshold, and calculate the objective function value for each random individual. If the objective function value of a random individual > the target threshold, then the corresponding random individual is regarded as a high-quality individual.
[0102] Specifically, the construction of the optimization population of the insertion tool is an experimental means to explore the characteristics of the first hot video clip, which can cover a wide range of possibilities and verify the impact of editing parameters on the quality of short videos. Quantify the performance of random individuals through the objective function to provide an experimental basis for subsequent optimization and production. The experimental results of random individuals can be used as the solution set of the optimization population for subsequent training of the editing optimization model to promote the intelligence of the video production process. Since the random population of the insertion tool does not rely on editing experience, it provides more objective support for the short video editing strategy.
[0103] In an embodiment of the present invention, constructing an optimization population includes:
[0104] Initialize and generate the optimization population based on the encoding of several high-quality individuals, specifically as follows:
[0105] Combine the writing style encoding, background music material encoding, and voice-over tone encoding of high-quality individuals respectively to obtain a writing style encoding library, a background music material encoding library, and a voice-over tone encoding library;
[0106] The writing style encoding of the optimization individuals in the optimization population is randomly selected from the writing style encoding library, the background music material encoding is randomly selected from the background music material encoding library, and the voice-over tone encoding is randomly selected from the voice-over tone encoding library to obtain several optimization individuals;
[0107] Calculate the objective function values of several optimization individuals based on the objective function, and regard the optimization individual with the largest objective function value as the target individual.
[0108] Specifically, in the previous stage of the present invention, by constructing a random population and evaluating it using the objective function, several high-quality individuals that meet specific conditions have been screened out. However, the random population inherently has a certain degree of randomness. Although it covers a wide range of possible spaces, it does not fully explore the optimal editing combinations. To further improve the quality of short video editing, an optimization-seeking population is constructed, aiming to find better target individuals based on high-quality individuals through specific optimization strategies on the existing high-quality combinations. In the experiment of the previous stage (random population), the random individuals with higher objective function values screened out were marked as high-quality individuals. Three key encodings are extracted from the high-quality individuals: the writing style encoding corresponds to the style of the connecting text, such as humorous, formal, narrative, etc. The background music material encoding corresponds to the type of background music. The voice-over tone encoding corresponds to the type of voice-over tone. The writing style encoding, background music material encoding, and voice-over tone encoding extracted from the high-quality individuals are respectively constructed into three independent encoding libraries. These encoding libraries cover the better parameter selections in the high-quality individuals, ensuring that the proven better combinations are preferentially considered in subsequent optimizations. Each optimization-seeking individual in the optimization-seeking population is randomly generated by combining from the three encoding libraries, so that the newly generated optimization-seeking population is based on the combination of high-quality individuals while retaining a certain degree of diversity. All optimization-seeking individuals are sorted by the objective function value, and the optimization-seeking individual with the largest objective function value is selected as the final target individual.
[0109] The target individual represents the optimal short video editing scheme for the first hot keyword.
[0110] In an embodiment of the present invention, an individual prediction model is trained, including:
[0111] Mapping the first hot keyword obtained within the first preset time period to a hot word vector, and merging the hot word vector and the corresponding target individual encoding to obtain a feature vector;
[0112] Inputting the feature vector into the individual prediction model, and the output of the individual prediction model represents the predicted evaluation score of the short video obtained by editing the first hot video corresponding to the first hot keyword according to the target individual encoding;
[0113] Taking the mean squared error between the predicted evaluation score and the objective function value corresponding to the target individual as the loss function of the individual prediction model, and updating the hyperparameters of the hidden layer of the individual prediction model through backpropagation.
[0114] Specifically, during the process of short video generation and optimization, the quality of the editing plan often needs to be measured by indicators such as the actual view count and interaction rate of the audience after the video is generated. This method is inefficient and has a high time cost. To improve efficiency, the present invention proposes an individual prediction model, which predicts the quality of a short video given a first hot keyword and a target individual encoding through training on historical data. This model can significantly reduce the trial-and-error cost and provide prediction support for the intelligent editing of short videos. The first hot keyword obtained within the first preset time period is converted into a numerical representation, i.e., a hot word vector, through a word embedding model (such as Word2Vec, BERT, or other word vector generation methods). The unstructured text data (keywords) is converted into numerical features suitable for processing by machine learning models. Each target individual (editing plan) consists of a writing style encoding, a background music material encoding, and a voice-over tone encoding. The various encodings of the target individual are merged in numerical form to form a complete target individual encoding. The hot word vector and the target individual encoding are concatenated to form a complete feature vector. This feature vector represents the combined characteristics of the first hot keyword and the target individual editing plan and is the input to the prediction model. The individual prediction model can be a deep learning-based model (MLP), which ensures that the model can accurately predict the actual quality of the short video corresponding to the target individual by minimizing the loss function. The backpropagation algorithm is used to update the model parameters. The optimizer can select common deep learning optimization algorithms (such as Adam, SGD, etc.). The weights and biases of the model's hidden layer are adjusted through the backpropagation algorithm, enabling the model to gradually converge to a better state.
[0115] In one embodiment of the present invention, the hidden layer of the individual prediction model is specifically as follows:
[0116] ;
[0117] ;
[0118] Wherein, represents the mapped output of the feature vector in the th hidden layer, represents the mapped output of the feature vector in the th hidden layer, represents the activation function, represents element-wise multiplication, represents based on constructed left diagonal matrix, the left diagonal matrix represents a sparse matrix with as the left diagonal element of the matrix and the remaining elements all being 0, represents based on constructed right diagonal matrix, the right diagonal matrix represents A sparse matrix with the right diagonal element being a matrix and the rest of the elements being 0, representing the feature space of the th hidden layer, representing the bias parameter of the feature space of the th hidden layer, representing the transpose operation.
[0119] Specifically, in the individual prediction model, the design of the hidden layer aims to map the input feature vector (including the word vector of the first hot keyword and the target individual encoding) into a high-dimensional feature space, extract key features, and then achieve accurate prediction of the short video clip quality. The hidden layer transforms the feature vector into a mapped output of the feature space through a series of matrix operations and non-linear activation functions, and dynamically adjusts the model parameters during the mapping process to optimize the prediction result. The input feature vector (formed by splicing the word vector of the first hot keyword and the genetic individual encoding) is mapped into a high-dimensional feature space after being processed by the hidden layer. The roles of the left diagonal matrix and the right diagonal matrix The introduction of the left diagonal matrix and the right diagonal matrix is to sparsify the mapping process, that is, to perform feature transformation only on specific dimensions and ignore features irrelevant to the prediction. This sparsification operation helps the model to more accurately extract features with a relatively high correlation with the first hot keyword and the genetic individual encoding. During the hidden layer mapping process, the model will focus on the relationship between the word vector of the first hot keyword and the target individual encoding, and then predict the evaluation score of the short video clip scheme. When the genetic individual of the input first hot keyword is significantly different from the target individual of the similar first hot keyword within the first preset time period, the mapping result of the feature vector will deviate from the high-quality region in the feature space. Since the model has learned the best matching relationship between the first hot keyword and the genetic individual through historical data during training, when deviating from this similarity, the model will give a lower prediction evaluation score. The mapping of the hidden layer has learned which feature combinations are of high quality. When deviating from the historical similar keywords, it is difficult for the model to determine the performance potential of the input scheme. However, when the genetic individual of the input first hot keyword is relatively close to the target individual of the similar first hot keyword in history, the model does not necessarily directly give a high prediction score. Although the similarity of the feature vector may increase the evaluation score, the model will also examine other features, such as subtle changes in the keyword, combinations of target individual encodings, and overall matching with the context. The model not only relies on feature similarity during the prediction process but also comprehensively considers the overall quality of the feature vector.
[0120] In an embodiment of the present invention, based on the genetic algorithm and the individual prediction model, the editing method of the second hot video is obtained, including:
[0121] Step 81, mapping the second hot keyword obtained in the second preset time period into a hot word vector ;
[0122] Step 82: Initialize a number of genetic individuals, which are represented as: ; where , and respectively represent the th writing style, the th music score material, and the th dubbing tone;
[0123] Step 83: Use the individual prediction model as the genetic adaptation function of the genetic population, and obtain the genetic adaptation function value of each genetic individual through the individual prediction model;
[0124] Step 84: Based on the genetic adaptation function values, sort each genetic individual from largest to smallest to obtain a feature ranking;
[0125] Step 85: Retain a preset number of genetic individuals in the feature ranking from front to back, and update the remaining individuals as the parent generation. The update includes:
[0126] Extract a random number X for each parent within a fixed value range. If X is greater than the first preset threshold P1, randomly select a parent for one-way crossover with the parent;
[0127] If X is less than the second preset threshold P2, mutate the parent;
[0128] If P2 < X < P1, pause the update of the parent;
[0129] Step 86: If steps 84 and 85 are looped for a preset number of times, stop the update, output the genetic individual with the largest genetic adaptation function value, obtain the editing method of the second hot video corresponding to the second hot keyword according to this genetic individual, and generate a corresponding short video.
[0130] Specifically, in the optimization of short video editing, since multiple variables (writing style, background music, dubbing, etc.) are involved, it is both inefficient and infeasible to directly find the optimal combination through the exhaustive method. Therefore, the present invention combines the genetic algorithm with the individual prediction model to simulate the selection and optimization process of biological evolution, and automatically generates a high-quality short video editing scheme based on the second hot keyword. The genetic algorithm uses the method of population evolution, and through operations such as selection, crossover, and mutation, gradually optimizes the initial genetic individuals (i.e., editing schemes), and combines the individual prediction model as the adaptation function to guide the optimization direction, and finally obtains the best-performing editing method.
[0131] A method and system for quickly batch producing short videos with a high completion rate, applied to the described method, includes:
[0132] A data acquisition module, configured to obtain the first hot keywords of an Internet platform at fixed time intervals within a first preset time period, and obtain the first hot videos corresponding to the first hot keywords;
[0133] A clip library module, configured to construct a video material library, which includes a text unit, a background music unit, and a voice-over unit; the text unit is used to generate text for the first hot video, the background music unit is used to generate background music for the first hot video, and the voice-over unit is used to generate a voice-over according to the generated text;
[0134] A population optimization module, configured to sequentially construct a random population and an optimization population based on the first hot video, and obtain a target individual according to the optimization population;
[0135] An individual prediction module, configured to obtain a number of target individuals and corresponding objective function values, combine the target individuals and the corresponding first hot keywords as training samples, and use the corresponding objective function values as sample labels to train an individual prediction model;
[0136] A clip planning module, within a second preset time period, obtains second hot keywords, and based on a genetic algorithm and the individual prediction model, obtains a clip method for the second hot video to generate a corresponding short video.
[0137] The embodiments of the present embodiment have been described, but the present embodiment is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of the present embodiment.
Claims
1. A method for quickly batch-forming short videos with a high completion rate, characterized in that: include: Step 1, within a first preset time period, obtaining a first hot keyword of an Internet platform at a fixed time interval, and obtaining a first hot video corresponding to the first hot keyword; Step 2, constructing a video material library, the video material library includes a text unit, a music unit and a dubbing unit; the text unit is used to generate text for the first hot video, the music unit is used to generate music for the first hot video, and the dubbing unit is used to generate dubbing according to the generated text; Step 3: construct a random population based on the first hot video to obtain a target individual of a high-quality combination, construct an optimal population on the target individual of the high-quality combination, and obtain a target individual based on the optimal population; Step 4, looping step 3 to obtain a number of target individuals and corresponding objective function values, combining the target individuals and the corresponding first hot keywords as training samples, and using the corresponding objective function values as sample labels to train and obtain an individual prediction model; Step 5, within the second preset time period, obtain the second hot keyword, and based on the genetic algorithm and the individual prediction model, obtain the editing method of the second hot video corresponding to the second hot keyword to generate the corresponding short video.
2. According to claim 1, a method for quickly batch-forming short videos with a high completion rate is characterized in that: Text unit, including video parsing tools and insertion tools, including: A video parsing tool, used to input the first hot video and extract the dialogue text of the first hot video; Insert tool, used to insert connecting text between dialogue texts. Connecting text includes: Step 21, obtaining the hot text corresponding to the first hot keyword, and disassembling the hot text based on the language model to obtain the event cause text, process text and result text of the hot text; Step 22, obtaining the timestamp corresponding to the dialogue text of the first hot video, if the time interval between adjacent dialogue texts is greater than a preset time, the interval between adjacent dialogue texts is used as an interval to be inserted; Step 23, the cause text, process text and result text, as well as the dialogue text inserted at both ends of the interval to be selected are mapped to word vectors respectively, and the correlation between the cause text, process text and result text and the interval to be selected is calculated based on the word vectors, and the calculation formula is as follows: ; in, Indicates the degree of association between the cause text, process text or result text and the interval to be inserted. Word vectors representing cause text, process text, or result text, The word vector representing the dialogue text inserted in front of the selected interval, The word vector representing the dialogue text inserted after the selected interval, , and represent the first weight parameter, the second weight parameter and the third weight parameter respectively, represents the Euclidean norm, represents the inner product. It means to find the square of Euclidean norm; Step 24, respectively determine the insertion candidate intervals with the maximum correlation of the cause text, the process text and the result text, and assign the cause text, the process text and the result text as the connection texts to the insertion candidate intervals with the maximum correlation.
3. The method for quickly batch-forming short videos with a high completion rate according to claim 1 is characterized in that: The dubbing unit and the music accompaniment unit are as follows: Obtain open source soundtrack materials and dubbing timbres, and cluster the soundtrack materials and dubbing timbres based on preset emotional styles to obtain several emotional soundtrack materials and several emotional dubbing timbres; Emotion recognition is performed on the first hot keyword to obtain the emotional features of the first hot keyword, and corresponding emotional music materials and emotional dubbing timbres are matched based on the emotional features.
4. A method for quickly batch-forming short videos with a high completion rate according to claim 2 or 3, characterized in that: Construct a random population, including: The open-source soundtrack materials and dubbing timbres are represented by real number codes; The style of the connecting text is converted through the language model, and the style is represented by real number coding; According to the first hot video, several random individuals that meet the constraints are initialized. The random individuals are encoded as follows: ;in, represents the i-th style of writing, represents the i-th soundtrack material, represents the i-th dubbing timbre; The constraints include: The emotional music material corresponding to the first hot video, The emotional dubbing tone corresponding to the first hot video; An objective function is constructed. The objective function is used to evaluate the short video obtained by randomly editing the first hot video. The formula of the objective function is as follows: ; ; ; in, represents the objective function value, represents the time weight, It represents the lifeline of short videos. , and Respectively represent the lifeline in , and A vector of time points, It represents the ratio of the number of likes to the number of views at the hth time point. It represents the ratio of the number of completed broadcasts to the number of broadcasts at the hth time point. represents the cross product of vectors, , and denote curvature weight, weighted playback weight and interaction weight respectively, , and are not 0, and , and The sum is 1, Indicates the preset end time point of the lifeline. Indicates the starting time point of the lifeline. represents the hth time point of the lifeline, The gth time point of the lifeline, Indicates the number of lifeline time points, The first index representing the number of lifeline time points, A second index indicating the number of lifeline time points, Not equal to , Represents the vector of the lifeline at the nth time point; Set the target threshold and calculate the objective function value for each random individual. If the objective function value of the random individual is greater than the target threshold, the corresponding random individual will be regarded as a high-quality individual.
5. The method for quickly batch-forming short videos with a high completion rate according to claim 4 is characterized in that: Constructing an optimal population, including: The optimal population is generated based on the encoding initialization of several high-quality individuals, as follows: The writing style codes, music material codes and dubbing timbre codes of high-quality individuals are combined respectively to obtain a writing style coding library, a music material coding library and a dubbing timbre coding library; The style codes of the optimal individuals of the optimal population are randomly selected from the style code library, the music material codes are randomly selected from the music material code library, and the dubbing timbre codes are randomly selected from the dubbing timbre code library, so as to obtain a number of optimal individuals; The objective function values of several optimizing individuals are calculated based on the objective function, and the optimizing individual with the largest objective function value is taken as the target individual.
6. The method for quickly batch-forming short videos with a high completion rate according to claim 5, characterized in that: The individual prediction model is trained, including: Mapping the first hot keyword obtained within the first preset time period into a hot word vector, and merging the hot word vector and the corresponding target individual code to obtain a feature vector; Inputting the feature vector into an individual prediction model, the output of the individual prediction model represents a prediction evaluation score of a short video obtained by editing the first hot video corresponding to the first hot keyword according to the target individual coding; The mean error variance of the predicted evaluation score and the objective function value corresponding to the target individual is used as the loss function of the individual prediction model, and the hyperparameters of the hidden layer of the individual prediction model are updated through back propagation.
7. The method for quickly batch-forming short videos with a high completion rate according to claim 6 is characterized in that: The hidden layers of the individual prediction model are as follows: ; ; in, Indicates that the eigenvector is The mapping output of the hidden layer is Indicates that the eigenvector is The mapping output of the hidden layer is express Activation function, represents element-wise multiplication, Indicates based on The left diagonal matrix constructed by As the left diagonal element of the matrix, the remaining elements are all 0 sparse matrix, Indicates based on The right diagonal matrix constructed by As a sparse matrix with the right diagonal elements of the matrix and the rest of the elements are 0, Indicates The feature space of hidden layers, Indicates The bias parameters of the feature space of the hidden layers, Represents a transpose operation.
8. The method for quickly batch-forming short videos with a high completion rate according to claim 7 is characterized in that: Based on the genetic algorithm and the individual prediction model, the editing method of the second hot video corresponding to the second hot keyword is obtained, including: Step 81: Map the second hot keyword obtained in the second preset time period into a hot word vector ; Step 82, initialize a number of genetic individuals, the genetic individuals are represented as: ;in, , and Respectively represent Type of writing style Music material and A kind of dubbing tone; Step 83, using the individual prediction model as the genetic fitness function of the genetic population, and obtaining the genetic fitness function value of each genetic individual through the individual prediction model; Step 84, based on the genetic fitness function value, sort each genetic individual from large to small to obtain a feature sorting; Step 85, retaining a preset number of genetic individuals in the feature sorting from front to back, and updating the remaining individuals as parents, the updating includes: For each parent generation, a random number X is drawn within a fixed value range. If X is greater than a first preset threshold value P1, a parent generation is randomly selected for one-way crossover; If X is less than the second preset threshold value P2, the parent generation is mutated; If P2<X<P1, then suspend updating the parent generation; Step 86, when step 84 and step 85 are looped for a preset number of times, the updating is stopped, and the genetic individual with the largest genetic fitness function value is output, and the editing method of the second hot video corresponding to the second hot keyword is obtained according to the genetic individual, and the corresponding short video is generated.
9. A system for quickly batch-forming short videos with a high completion rate, characterized in that: The method according to any one of claims 1 to 8, comprising: A data collection module, used to obtain a first hot keyword on the Internet platform at a fixed time interval within a first preset time period, and obtain a first hot video corresponding to the first hot keyword; A clip library module, used to construct a video material library, the video material library includes a text unit, a music unit and a dubbing unit; the text unit is used to generate text for the first hot video, the music unit is used to generate music for the first hot video, and the dubbing unit is used to generate dubbing according to the generated text; A population optimization module, used to construct a random population based on the first hot video to obtain a target individual of a high-quality combination, construct an optimal population on the target individual of the high-quality combination, and obtain a target individual according to the optimal population; An individual prediction module is used to obtain a number of target individuals and corresponding objective function values, and to use the target individuals and the corresponding first hot keywords as training samples, and the corresponding objective function values as sample labels, so as to train and obtain an individual prediction model; The editing planning module obtains the second hot keyword within a second preset time period, and obtains the editing method of the second hot video based on a genetic algorithm and an individual prediction model to generate a corresponding short video.
Citation Information
Patent Citations
AIGC method for intelligently generating short video through text
CN116977903A
Generative model optimization method and device, electronic equipment and storage medium
CN118709746A