Video generation method and device, electronic equipment and computer readable storage medium
By optimizing the sentences to be processed, ensuring the consistency of the video picture and video speech duration, the problem of reducing the correlation between materials and copywriting sentences in the prior art is solved, and the effect of video generation is improved.
Patent Information
- Application Number
- CN202411886210.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, when generating videos, the language time corresponding to the copy sentence is greater than the length of the material, resulting in the worse the correlation between the material and the copy sentence as the timeline goes back, and the poorer the generation effect or failure.
By optimizing the sentences to be processed, the sentence length is reduced, and the video screen duration is consistent with the video speech duration, thereby improving the relevance of the material and copycat sentences when generating videos.
By optimizing the sentences to be processed, the correlation between the material and copywriting sentences in the video timeline is avoided, which improves the overall effect of video generation and avoids generation failure.
Smart Images

Figure CN120050479A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of image and video processing, and in particular, to technical fields such as artificial intelligence, large models, video synthesis, etc. Specifically, the present disclosure relates to a video generation method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] As a kind of information medium, videos are becoming increasingly popular among people because of their rich media form, which can provide an immersive experience, and more and more people are starting to participate in video creation.
[0003] Compared with simple and fast text creation, video creation requires a lot of energy, but the readability of text creation is significantly lower than that of video creation. The creation method of generating videos based on text does not require a lot of energy for creation and can also produce videos with stronger dissemination. Summary of the Invention
[0004] The present disclosure provides a video generation method and apparatus, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect of the present disclosure, there is provided a video generation method, the method comprising:
[0006] Obtaining the text information to be processed, and segmenting the text information to be processed to obtain a plurality of sentences to be processed;
[0007] Based on the sentences to be processed, obtaining at least one candidate material that matches the sentences to be processed, and synthesizing the voice to be processed corresponding to the sentences to be processed;
[0008] In the case where the duration of the video frame is less than the duration of the video voice, optimizing the sentences to be processed, and generating a target video based on the optimized sentences to be processed and the candidate materials;
[0009] Wherein, the duration of the video frame is the sum of the time lengths corresponding to all the candidate materials, and the duration of the video voice is the time length corresponding to the voice to be processed.
[0010] According to a second aspect of the present disclosure, there is provided a video generation apparatus, the apparatus comprising:
[0011] A preprocessing module, configured to obtain the text information to be processed, and segment the text information to be processed to obtain a plurality of sentences to be processed;
[0012] A material obtaining module, configured to obtain at least one candidate material that matches the sentences to be processed based on the sentences to be processed, and synthesize the voice to be processed corresponding to the sentences to be processed;
[0013] A material optimization module, configured to optimize the sentence to be processed when the duration of the video picture is less than the duration of the video voice, and generate a target video based on the optimized sentence to be processed and the candidate materials;
[0014] Wherein, the duration of the video picture is the sum of the time lengths corresponding to all the candidate materials, and the duration of the video voice is the time length corresponding to the voice to be processed.
[0015] According to a third aspect of the present disclosure, there is provided an electronic device, which includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above video generation method.
[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above video generation method.
[0020] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the above video generation method when executed by a processor.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a flowchart of a video generation method provided by an embodiment of the present disclosure;
[0024] Figure 2 is a flowchart of partial steps of another video generation method provided by an embodiment of the present disclosure;
[0025] Figure 3 is a flowchart of partial steps of another video generation method provided by an embodiment of the present disclosure;
[0026] Figure 4 is a flowchart of partial steps of another video generation method provided by an embodiment of the present disclosure;
[0027] Figure 5 It is a schematic flowchart of some steps of another video generation method provided by an embodiment of the present disclosure;
[0028] Figure 6 It is a schematic diagram of the process of a specific embodiment of the video generation method provided by an embodiment of the present disclosure;
[0029] Figure 7 It is a schematic structural diagram of a video generation device provided by an embodiment of the present disclosure;
[0030] Figure 8 It is a block diagram of an electronic device for implementing the video generation method of an embodiment of the present disclosure. Specific Embodiments
[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0032] In some related technologies, based on the understanding of the copywriting, the shot language is generated, and then the material library is retrieved to obtain materials. The text-video is paired using the scheduling algorithm to form a timeline.
[0033] In specific implementation, a dual-tower multi-modal model can be used to calculate the relevance between the materials in the material library and each copywriting sentence. The material with the highest relevance is selected for each copywriting sentence to be paired to generate a timeline.
[0034] However, during the generation process of the timeline, the situation of insufficient materials may occur, that is, the language duration corresponding to the copywriting sentence is longer than the duration of the materials. For example, there may be only one screenshot for some events, so that the later the timeline is, the worse the relevance between the materials and the copywriting sentences, resulting in a poor overall generation effect or generation failure.
[0035] The video generation method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.
[0036] The video generation method provided by the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in a memory. Alternatively, the method can be executed by a server.
[0037] Figure 1 The flowchart of the video generation method provided by the embodiments of the present disclosure is shown. As Figure 1 shown, the video generation method provided by the embodiments of the present disclosure may include step S110, step S120, and step S130.
[0038] In step S110, the text information to be processed is obtained, and the text information to be processed is segmented to obtain a plurality of sentences to be processed;
[0039] In step S120, based on the sentences to be processed, at least one candidate material matching the sentences to be processed is obtained, and the speech to be processed corresponding to the sentences to be processed is synthesized;
[0040] In step S130, when the duration of the video frame is less than the duration of the video speech, the sentences to be processed are optimized, and based on the optimized sentences to be processed and the candidate materials, a target video is generated;
[0041] Among them, the duration of the video frame is the sum of the time lengths corresponding to all candidate materials, and the duration of the video speech is the time length corresponding to the speech to be processed.
[0042] For example, in step S110, the text information to be processed may be a video copywriting uploaded by a user or input by a user through an interaction device, which may be one or several paragraphs, and each paragraph includes one or more complete sentences.
[0043] In some possible implementation manners, segmenting the text information to be processed may be based on terminators in the video copywriting, such as punctuation information (such as a period, a comma, a semicolon, etc.), and using the punctuation as a dividing boundary to segment the video copywriting into multiple sentences, and each sentence is a sentence to be processed.
[0044] In some possible implementation manners, in step S120, for each sentence to be processed, materials matching the sentence to be processed are obtained from a material library as candidate materials, and the speech to be processed corresponding to the sentence to be processed is synthesized based on the TTS (Text To Speech) technology.
[0045] Among them, the candidate materials can be video frame materials or picture materials. The video frame materials and picture materials can be obtained separately from different material libraries as candidate materials.
[0046] In some possible implementation manners, in step S130, a plurality of candidate materials are combined into a video frame in sequence, and the duration of the video frame is the time length corresponding to the video frame composed of the plurality of candidate materials. The duration of the video voice is the duration of the voice to be processed.
[0047] Specifically, the staying time of the picture material in the video frame is set to a preset staying time, such as 5 seconds, and the video frame material stays in the video frame according to the actual duration of the video frame material. Therefore, the time length corresponding to the picture material is the preset staying time, and the time length corresponding to the video frame material is the actual duration of the video frame material.
[0048] In some possible implementation manners, the duration of the video frame is compared with the duration of the video voice. When the duration of the video frame is the same as the duration of the video voice, the video frame and the voice to be processed are combined into a video, and the text content corresponding to the sentence to be processed is added as subtitles to the video as the target video corresponding to the sentence to be processed.
[0049] When the duration of the video frame is greater than the duration of the video voice, the time length corresponding to the video frame can be reduced by editing the video frame material or adjusting the staying time of the picture material. The processed video frame and the voice to be processed are combined into a video, and the text content corresponding to the sentence to be processed is added as subtitles to the video as the target video corresponding to the sentence to be processed.
[0050] When the duration of the video frame is less than the duration of the video voice, the sentence to be processed can be optimized to reduce the length of the sentence to be processed, and further reduce the length of the voice to be processed corresponding to the optimized sentence to be processed. The video frame and the voice to be processed corresponding to the optimized sentence to be processed are combined into a video, and the text content corresponding to the optimized sentence to be processed is added as subtitles to the video as the target video corresponding to the sentence to be processed.
[0051] After obtaining the target videos corresponding to all the sentences to be processed, the videos are arranged in the order of the sentences to be processed in the text information to be processed to form a time axis, and video rendering is performed to generate a video corresponding to the text information to be processed.
[0052] In the video generation method provided by the embodiments of the present disclosure, when the duration of the video frame is less than the duration of the video voice, that is, when the candidate materials corresponding to the sentence to be processed are insufficient, the sentence to be processed is optimized to make the video frame correspond to the video voice, avoiding the worse the correlation between the video frame and the video voice or the video copywriting as the video timeline goes further, improving the overall video generation effect and avoiding video generation failure.
[0053] The following specifically introduces the video generation method provided by the embodiments of the present disclosure.
[0054] Figure 2 The flowchart shows an implementation manner of obtaining at least one candidate material that matches the sentence to be processed based on the sentence to be processed, as Figure 2 shown, which may include step S210, step S220, step S230, and step S240.
[0055] In step S210, keywords are extracted from the sentence to be processed, and multiple retrieval materials that match the sentence to be processed are obtained based on the obtained keywords;
[0056] In step S220, a multi-modal large model is used to calculate the first correlation score between the retrieval material and the sentence to be processed;
[0057] In step S230, a multi-modal large model is used to generate a copywriting based on the retrieval material, obtain the video copywriting corresponding to the retrieval material, and calculate the second correlation score between the video copywriting and the sentence to be processed;
[0058] In step S240, the first correlation score and the second correlation score are weighted and summed. When the weighted sum result is greater than a preset threshold, the retrieval material is determined as a candidate material.
[0059] In some possible implementation manners, in step S210, keyword extraction technology is used to extract one or more keywords from the sentence to be processed as a Query (query term), and the Query is used to retrieve in the material library, and the retrieval result is used as the retrieval material that matches the sentence to be processed.
[0060] In some possible implementation manners, the candidate material can be a video frame material or a picture material. That is to say, the retrieval material can be a video frame material or a picture material.
[0061] In some possible implementation manners, the Query can be used to retrieve in different material libraries to respectively obtain the video frame material and the picture material as the retrieval material.
[0062] That is, based on the obtained keywords (i.e., Query), at least one video frame material that matches the sentence to be processed is obtained from the video material library as the retrieval material; based on the obtained keywords (i.e., Query), at least one picture material that matches the sentence to be processed is obtained from the picture material library as the retrieval material.
[0063] Among them, the video material library includes multiple video frame materials, which can be a pre-generated video database (such as a database composed of Internet data). The picture material library includes multiple pictures, which can be a pre-generated picture database. Obtaining retrieval materials from the video material library and the picture material library based on Query can utilize the technology of retrieving in the database based on Query, which will not be elaborated here.
[0064] The quantities of the video frame materials and picture materials as the retrieval materials can be preset to avoid excessive quantities of the obtained retrieval materials, resulting in too much data to be processed and affecting the efficiency of the entire video generation.
[0065] That is, in the retrieval results based on Query, the preset first quantity of video frame materials with the highest relevance to Query is obtained as the retrieval materials, and in the retrieval results based on Query, the preset second quantity of picture materials with the highest relevance to Query is obtained as the retrieval materials.
[0066] Among them, the first quantity and the second quantity can be the same or different.
[0067] In some possible implementation manners, in step S220, for each retrieval material, a multimodal large model is used to calculate the relevance between the retrieval material and the sentence to be processed, and a first relevance score is obtained.
[0068] Among them, the multimodal large model used can be a two-tower model, which has two encoders. One encoder is used to extract the features of the video frame material or picture material, and the other encoder is used to extract the features of the sentence to be processed. The relevance between the retrieval material and the sentence to be processed is calculated based on the extracted features. The higher the relevance between the retrieval material and the sentence to be processed, the higher the obtained first relevance score.
[0069] In some possible implementation manners, in step S230, for each retrieval material, a multimodal large model is used to generate a descriptive copy for the retrieval material based on the retrieval material, that is, the video copy corresponding to the retrieval material. By calculating the relevance between the video copy and the sentence to be processed, a second relevance score is obtained.
[0070] Since the video copywriting is a descriptive copy of the retrieval materials, the higher the similarity between the video copywriting and the sentence to be processed, the higher the relevance of the description of the retrieval materials to the sentence to be processed, the higher the relevance of the retrieval materials to the sentence to be processed, and the higher the second relevance score obtained.
[0071] Among them, the multimodal large model used can be the DeepSeek model (multimodal description model). Calculating the similarity between the video copywriting and the sentence to be processed can be based on the large language model to calculate the similarity between the video copywriting and the sentence to be processed.
[0072] In some possible implementation manners, in step S240, for each retrieval material, a weighted sum of the first relevance score and the second relevance score corresponding to the retrieval material is calculated, that is, the first relevance score and the second relevance score are respectively multiplied by their corresponding weighted weights, and the multiplication results are added to obtain the matching score of the retrieval material and the sentence to be processed.
[0073] The higher the matching score of the retrieval material and the sentence to be processed, the higher the relevance of the retrieval material and the sentence to be processed. Therefore, when the matching score corresponding to the retrieval material is greater than the preset threshold, it indicates that the relevance of the retrieval material and the sentence to be processed meets the requirements, and the retrieval material can be used as a candidate material for the sentence to be processed.
[0074] Through the above method, the obtained retrieval materials can be screened from multiple aspects, thereby improving the relevance between the obtained candidate materials and the sentence to be processed.
[0075] As described above, in some possible implementation manners, after obtaining the candidate materials, the time lengths corresponding to all the candidate materials are added up to obtain the video screen duration.
[0076] Among them, the time length corresponding to the picture material is the preset stay time, and the time length corresponding to the video screen material is the actual duration of the video screen material.
[0077] When the video screen duration is less than the video voice duration (the acquisition method is as described above and will not be elaborated here), by optimizing the sentence to be processed to reduce the length of the sentence to be processed, and then reducing the length of the to-be-processed voice corresponding to the optimized sentence to be processed, the video is composed of the video screen and the to-be-processed voice corresponding to the optimized sentence to be processed, and the text content corresponding to the optimized sentence to be processed is added as subtitles to the video as the video corresponding to the sentence to be processed.
[0078] Figure 3 Shows a schematic flowchart of an implementation manner of optimizing the sentence to be processed and generating a target video based on the optimized sentence to be processed and candidate materials. As Figure 3As shown, optimizing the sentence to be processed and generating a target video based on the optimized sentence to be processed and candidate materials may include step S310, step S320, and step S330.
[0079] In step S310, count the number of reference words corresponding to the speech with a time length consistent with the video frame duration.
[0080] In step S320, optimize the sentence to be processed so that the difference between the number of words contained in the optimized sentence to be processed and the number of reference words is within a preset range.
[0081] In step S330, based on the optimized sentence to be processed, generate the video speech corresponding to the optimized sentence to be processed, form the video frames based on the candidate materials, and generate the target video based on the video speech and the video frames.
[0082] In some possible implementation manners, in step S310, through statistics, obtain the number of words contained in the speech with a preset speech rate and a preset speech intonation within the time with a time length consistent with the video frame duration, that is, the number of reference words.
[0083] In some possible implementation manners, in step S320, optimize the sentence to be processed based on a pre-trained large model so that the difference between the number of words contained in the optimized sentence to be processed and the number of reference words is within a preset range.
[0084] Among them, the large model may be a multimodal large model, which is used to rewrite the sentence to be processed, and rewrite the sentence to be processed into a sentence with a meaning close to that of the sentence to be processed but with a different number of words contained.
[0085] In some possible implementation manners, in step S330, synthesize the speech corresponding to the optimized sentence to be processed based on the TTS technology, that is, the video speech, form a video by combining the video frames composed of the candidate materials and the video speech, and add the text content corresponding to the optimized sentence to be processed as subtitles into the video as the target video corresponding to the sentence to be processed.
[0086] The video frames composed of the candidate materials may be the video frames composed of the candidate materials in sequence. Among them, for the picture materials, set the residence time in the video frames to a preset residence time. That is to say, for each picture material, generate a video frame with a preset residence time, and for each video frame material, directly use the material as the video frame, and form the final video frame by combining the video frames corresponding to all the candidate materials in sequence.
[0087] As described above, in step S320, the sentence to be processed can be optimized based on a pre-trained large model, so that the difference between the number of words included in the optimized sentence to be processed and the number of reference words is within a preset range.
[0088] Of course, the sentence to be processed can also be optimized based on other machine learning models, and any method that can optimize the sentence to be processed is within the protection scope of the embodiments of the present disclosure.
[0089] Figure 4 The flowchart shows an implementation manner of optimizing the sentence to be processed based on a pre-trained large model, as Figure 4 shown, optimizing the sentence to be processed based on a pre-trained large model may include step S410 and step S420.
[0090] In step S410, determine the optimized word number range based on the reference word number;
[0091] In step S420, at least use the sentence to be processed and the optimized word number range to construct a prompt text, and input the prompt text into the pre-trained large model to obtain an optimized sentence to be processed whose included word number is within the optimized word number range.
[0092] In some possible implementation manners, in step S410, subtract a predetermined value from the reference word number as the upper bound of the optimized word number range, and add a predetermined value to the reference word number as the lower bound of the optimized word number range to generate the optimized word number range.
[0093] Through such a generation method, it can be ensured that the reference word number is within the optimized word number range, and the differences between the upper and lower bounds of the word number range and the reference word number are both within the preset range (i.e., within the predetermined value range).
[0094] In some possible implementation manners, in step S420, the pre-trained large model can be an LLM.
[0095] LLMs are deep learning models trained using large amounts of text data, which can generate natural language text or understand the meaning of language text. They are characterized by their large scale and huge number of parameters (usually reaching over tens of billions), and are typically based on deep learning architectures such as the Transformer architecture. The difference between LLMs and ordinary pre-trained language models lies in the parameter scale. When the parameter scale exceeds a certain level, the model achieves a significant performance improvement and exhibits capabilities that small models do not have, such as in-context learning ability, being able to learn complex patterns in language and perform a wide range of tasks, including text summarization, modification, translation, sentiment analysis, multi-turn conversations, and so on. Generally speaking, a language model based on a deep learning architecture with a parameter scale of over tens of billions can be considered a large language model.
[0096] Common LLMs include: GPT-3 (Generative Pre-trained Transformer 3), T5 (Text-to-Text Transfer Transformer), GPT-4, PaLM (a large language model proposed by Google), LLaMA (Large Language Model Meta AI, a large language model released by Meta AI), and so on.
[0097] In large language models, the prompt text is the input text used to guide the model to generate specific types of text or perform specific tasks. Among them, the prompt text can be generated in various ways. It can be the text obtained by directly splicing the natural language question text and the candidate key information as the prompt text, or predefined templates or rules can be used to generate the prompt text. Among them, the prompt template describes the structure and content of the required prompt text.
[0098] The embodiments of the present disclosure can utilize the text modification function of the large model to optimize the sentence to be processed.
[0099] Specifically, the sentence to be processed and the range of the number of optimized words can be filled into a preset prompt template to obtain the prompt text. The prompt template can include: an instruction to prompt the large model to rewrite the sentence to be processed so that the rewritten sentence to be processed meets the range of the number of optimized words.
[0100] For example, the prompt template can be: You are a professional video copywriter. Rewrite {} into a video description within the range of {} words. The generated prompt text can be: You are a professional video copywriter. Rewrite {sentence to be processed} into a video description within the range of {range of the number of optimized words} words.
[0101] The sentence to be processed is a sentence in the text information to be processed, which has context information, that is, the sentences adjacent to the sentence to be processed in the text information to be processed. Using the context learning ability of the large model, information that can help rewrite the sentence to be processed can also be obtained from the context above and below the sentence to be processed.
[0102] Therefore, the context above and below the sentence to be processed in the text information to be processed can also be used as the content of the prompt text and input into the large model, so that the large model can obtain information from the context of the sentence to be processed to help rewrite the sentence to be processed.
[0103] Similarly, the sentence to be processed, the optimized word count range, and the context of the sentence to be processed can be filled into a preset prompt template to obtain the prompt text. Among them, the prompt template can include: prompting the large model to obtain information from the context of the sentence to be processed and rewrite the sentence to be processed so that the rewritten sentence to be processed meets the instruction of the optimized word count range.
[0104] For example, the prompt template can be: You are a professional video copywriter. Please rewrite {} into a video description within the word count range of {text1 - text2}, and closely combine the context: the context above {}, the context below {}, and try to ensure that the copy does not deviate from the theme and the context logic is smooth.
[0105] The generated prompt text can be: You are a professional video copywriter. Please rewrite {the sentence to be processed} into a video description within the word count range of {the optimized word count range}, and closely combine the context: the context above {the context above the sentence to be processed in the text information to be processed}, the context below {the context below the sentence to be processed in the text information to be processed}, and try to ensure that the copy does not deviate from the theme and the context logic is smooth.
[0106] In some specific implementation manners, since the sentence to be processed is a part of the text information to be processed, and the text information to be processed is used as a video copy, which has a theme, therefore, the sentence to be processed is also a sentence centered around the theme, and the optimized sentence to be processed should also be centered around the theme. Therefore, the theme can also be used as a part of the prompt text to provide information for the large model.
[0107] That is to say, the sentence to be processed, the theme, the optimized word count range, and the context of the sentence to be processed can be filled into a preset prompt template to obtain the prompt text. Among them, the prompt template can include: prompting the large model to obtain information from the context and theme of the sentence to be processed and rewrite the sentence to be processed so that the rewritten sentence to be processed meets the instruction of the optimized word count range.
[0108] For example, the prompt template can be: You are a professional video copywriter. Please rewrite the following sentence for {}, creating a video description within the word count range of {} based on the theme {}. Please closely integrate the context: the above text {}, the following text {}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0109] The generated prompt text can be: You are a professional video copywriter. Please rewrite the following sentence for {sentence to be processed}, creating a video description within the word count range of {} based on the theme {theme}. Please closely integrate the context: the above text {previous text adjacent to the sentence to be processed in the text information to be processed}, the following text {next text adjacent to the sentence to be processed in the text information to be processed}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0110] In some specific implementation manners, since the sentence to be processed can be regarded as a description of the candidate material, therefore, the candidate material can provide information for rewriting the sentence to be processed. However, since the candidate material and the sentence to be processed are information in different modalities, it is necessary to use a multimodal large model to utilize the information provided by the candidate material.
[0111] Multimodal large models, namely Multimodal Large Language Models (MLLM) or vision multimodal models, namely VLM (Vision Language Models), are based on Large Language Models (LLM) and Large Vision Models (LVM). They can process various media data types including text, images, audio, and video, learn the associations between data of different modalities through joint training, and improve the performance and generalization ability of the model. The core lies in cross-modal information fusion and understanding, enabling the model to more comprehensively and accurately grasp the deep meaning behind the data.
[0112] Therefore, through the multimodal large model, it is possible to achieve information fusion and understanding of the sentence to be processed and the candidate material, so as to grasp the deep meaning behind the sentence to be processed and the candidate material, and further improve the rewriting effect of the large model on the sentence to be processed.
[0113] Specifically, the sentence to be processed, the theme, the word count range for optimization, the context of the sentence to be processed, and the video frame generated based on the candidate material can be filled into a preset prompt template to obtain the Prompt (hint) input to the multimodal model. The prompt template can include: prompting the multimodal large model to obtain information from the context, theme, and candidate material of the sentence to be processed, and rewrite the sentence to be processed so that the rewritten sentence to be processed meets the instruction of the word count range for optimization.
[0114] For example, the prompt template can be: You are a professional video copywriter. Please rewrite the current video {} for {}, into a video description based on the theme {} within the number of words {}, in close combination with the context: the above text {}, the below text {}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0115] The generated prompt text can be: You are a professional video copywriter. Please rewrite the current video {video footage generated according to candidate materials} for {sentence to be processed}, into a video description based on the theme {theme} within the optimized word count range {}, in close combination with the context: the above text {preceding text adjacent to the sentence to be processed in the text information to be processed}, the below text {following text adjacent to the sentence to be processed in the text information to be processed}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0116] Whether using LLM, MLLM, or VLM to optimize the sentence to be processed first, it is necessary to pre-train the large model used. Figure 5 The flowchart shows an implementation method for training the large model, as Figure 5 shown, pre-training the large model used may include step S510, step S520, and step S530.
[0117] In step S510, obtain multiple training videos and the training texts corresponding to the training videos, segment the training texts, and obtain multiple ground-truth sentences.
[0118] In step S520, rewrite the ground-truth sentences into training sentences with a different number of words from the ground-truth sentences.
[0119] In step S530, construct a training prompt text using at least the training sentences and the training word count range, input the training prompt text into the large model, and train the large model according to the correlation between the output of the large model and the ground-truth sentences.
[0120] Among them, the number of words in the ground-truth sentences is within the training word count range; the differences between the upper and lower bounds of the training word count range and the number of words in the ground-truth sentences are both within the preset range.
[0121] In some possible implementation manners, in step S510, obtain video-text pairs with relatively good quality from the material library or online, use the videos in the video-text pairs as training videos, and use the texts in the video-text pairs as the training texts corresponding to the training videos.
[0122] By obtaining multiple video-text pairs, obtain multiple training videos and the training texts corresponding to the training videos.
[0123] The method of splitting the training text to obtain multiple true-value sentences can be the same as the method of splitting the text information to be processed to obtain multiple sentences to be processed, which will not be elaborated here.
[0124] Based on the obtained true-value sentences, the training video is also split to obtain the training video frames of the training video corresponding to each true-value sentence. Specifically, the video frame in the training video where the subtitle corresponding to the true-value sentence appears can be used as the training video frame corresponding to the true-value sentence.
[0125] In some possible implementation manners, in step S520, an LLM such as GPT can be used to rewrite the true-value sentence, and the true-value sentence is rewritten into a training sentence with a different number of words but a similar meaning to that contained in the true-value sentence.
[0126] For one true-value sentence, multiple corresponding training sentences can be generated through rewriting.
[0127] Similar to the reference number of words and the optimized range of the number of words, the number of words contained in the true-value sentence is within the training range of the number of words; the differences between the upper and lower bounds of the training range of the number of words and the number of words contained in the true-value sentence are both within the preset range.
[0128] In some possible implementation manners, in step S530, the training sentence and the training range of the number of words can be filled into a preset prompt template to obtain a prompt text. The prompt template may include: an instruction for prompting the large model to rewrite the training sentence so that the rewritten training sentence meets the training range of the number of words.
[0129] For example, the prompt template can be: You are a professional video copywriter. For {}, rewrite it into a video description within the range of {} words. The generated prompt text can be: You are a professional video copywriter. For {training sentence}, rewrite it into a video description within the range of {training range of the number of words} words.
[0130] Input the prompt text into the large model to be trained to obtain the output of the large model. Take the function of calculating the correlation between the output of the large model and the true-value sentence as the loss function of the large model, take the correlation between the output of the large model and the true-value sentence as the loss function value, and modify the parameters of the large model by backpropagating the loss function value to implement the training of the large model.
[0131] In the case where the large model to be trained is a multimodal large model, just as the sentence to be processed is a sentence in the text information to be processed, it has context information, the text information to be processed has a theme, the sentence to be processed can be regarded as a description of the candidate material, and the information that can help rewrite the sentence to be processed includes the above and below texts of the sentence to be processed, the main body of the text information to be processed, and the video frame of the candidate material. The true value sentence is a sentence in the training text, and the above and below texts adjacent to the true value sentence in the training text, the training video frame of the training video corresponding to the true value sentence, and the theme of the training text can also provide information for the training of the large model.
[0132] Therefore, the training sentence, the theme of the training text, the range of the number of training words, the above and below texts adjacent to the true value sentence in the training text, and the training video frame of the training video corresponding to the true value sentence can be filled into a preset prompt template to obtain the Prompt (hint) input into the multimodal model. Among them, the prompt template can include: instructing the multimodal large model to obtain information from the above and below texts adjacent to the true value sentence in the training text, the theme of the training text, and the training video frame of the training video corresponding to the true value sentence, and rewrite the training sentence so that the rewritten training sentence meets the range of the number of training words.
[0133] For example, the prompt template can be: You are a professional video copywriter. Please, according to the current video {}, for {}, rewrite it into a video description based on the theme {} within the range of {} words. Please closely combine the context: the above text {}, the below text {}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0134] The generated prompt text can be: You are a professional video copywriter. Please, according to the current video {the training video frame of the training video corresponding to the true value sentence}, for {the training sentence}, rewrite it into a video description based on the theme {the theme of the training text} within the range of {the range of the number of training words} words. Please closely combine the context: the above text {the above text adjacent to the true value sentence in the training text}, the below text {the below text adjacent to the true value sentence in the training text}, and try to ensure that the copy does not deviate from the theme and the context is logically coherent.
[0135] Input the prompt text into the large model to be trained to obtain the output of the large model. Calculate the correlation function between the output of the large model and the true sentence, the consistency function between the output of the large model and the context before and after the true sentence in the training text, and the correlation function between the output of the large model and the training video frame as the loss function of the large model. Take the weighted sum of the correlation between the output of the large model and the true sentence, the consistency between the output of the large model and the context before and after the true sentence in the training text, and the correlation between the output of the large model and the training video frame as the loss function value. Modify the parameters of the large model by backpropagating the loss function value to train the large model.
[0136] Among them, an LLM such as GPT can be used to calculate the consistency between the output of the large model and the context before and after the true sentence in the training text, and an MLLM such as gpt4o (the language model released by OpenAI for the chatbot ChatGPT) can be used to calculate the correlation between the output of the large model and the training video frame.
[0137] After optimizing the sentence to be processed, synthesize the voice corresponding to the optimized sentence to be processed based on the TTS technology, that is, the video voice. Combine the video frame composed of candidate materials and the video voice to form a video, and add the text content corresponding to the optimized sentence to be processed as subtitles into the video as the video corresponding to the sentence to be processed.
[0138] As described above, in some possible implementation manners, the duration of the video frame is greater than the duration of the video voice. The duration corresponding to the video frame can be reduced by editing the video frame materials or adjusting the staying time of the picture materials.
[0139] When the duration of the video frame is greater than the duration of the video voice and the difference between the duration of the video frame and the duration of the video voice is not greater than the preset duration threshold, modify the time length corresponding to the picture materials in the candidate materials so that the modified duration of the video frame is equal to the duration of the video voice.
[0140] That is to say, the difference between the duration of the video frame and the duration of the video language is small, and the consistency between the modified duration of the video frame and the duration of the video voice can be achieved by reducing the staying time of the picture materials in the video frame.
[0141] Therefore, only need to modify the staying time of the picture materials in the video frame, regenerate the video frame, combine the video frame with the voice to be processed to form a video, and add the text content corresponding to the sentence to be processed as subtitles into the video as the video corresponding to the sentence to be processed.
[0142] When the duration of the video picture is greater than the duration of the video voice, and the difference between the duration of the video picture and the duration of the video voice is greater than a preset duration threshold, crop the video picture material in the candidate material so that the duration of the cropped video picture is equal to the duration of the video voice.
[0143] That is to say, the difference between the duration of the video picture and the duration of the video language is relatively large. Reducing the duration of the picture material staying in the video picture to make the duration of the modified video picture consistent with the duration of the video voice may result in the picture material staying in the video picture for a short time, which may cause the user to ignore the picture material. Cropping the video picture material has less information loss for the overall expression of the video picture, and the user can still obtain the information expressed by the video picture material. Therefore, it is necessary to crop the video picture material in the candidate material to make the duration of the modified video picture consistent with the duration of the video voice.
[0144] Regenerate the video picture after cropping, combine the video picture with the voice to be processed to form a video, and add the text content corresponding to the sentence to be processed as subtitles into the video as the video corresponding to the sentence to be processed.
[0145] After obtaining the videos corresponding to all the sentences to be processed by the above method, arrange the videos in the order of the sentences to be processed in the text information to be processed to form a timeline, and perform video rendering to generate the video corresponding to the text information to be processed.
[0146] The following uses a specific embodiment to introduce the video generation method provided by the embodiments of the present disclosure in detail.
[0147] Figure 6 Shows a schematic process diagram of a specific embodiment of the video generation method provided by the embodiments of the present disclosure. As Figure 6 shown, input the video copywriting, split the video copywriting according to terminators, etc. into different sentences, and at the same time perform Query extraction on the split copywriting sentences. For each Query, perform material retrieval in the video picture material library to obtain 20 video picture materials, perform picture material retrieval across the network to obtain 10 picture materials, and use the obtained video picture materials and picture materials as retrieval materials. Use a multi-modal relevance model to score the relevance of 30 retrieval materials and the copywriting sentences, use the DeepSeek multi-modal description model to generate copywriting for 30 retrieval materials, and at the same time calculate the relevance score of the generated copywriting and the copywriting sentences, and perform weighted summation of the above two scores as the final video-copywriting matching score. Filter the retrieval materials based on the score threshold, and select the retrieval materials greater than the threshold as candidate materials suitable for the copywriting sentence currently being processed.
[0148] For the picture materials in the candidate materials, set the residence duration of them in the generated video screen to 5s. For the duration of the video screen, set the residence duration of it in the generated video screen to the actual duration of the video screen. The sum of the residence durations of all candidate materials is the available duration tv of the candidate materials.
[0149] The current copywriting sentence being processed is pre-generated by TTS, and it is determined that the required video screen duration of the current copywriting sentence being processed is tt.
[0150] Compare the durations of the two. If tv - tt > 3, then crop the video screen materials in the time dimension to make tt the same as tv. If 0 <= tv - tt < 3, then change the residence duration of the picture materials to 5 - (tv - tt) to ensure that the picture residence duration is within a certain range and will not be too fast or too slow, affecting the experience.
[0151] If tv - tt < 0, then optimize the current copywriting sentence being processed. Optimize the current copywriting sentence being processed based on the candidate materials, the current copywriting sentence being processed, the video copywriting theme, the sentences adjacent to the current copywriting sentence in the video copywriting, etc.
[0152] First, based on statistical methods, count the number of words corresponding to tv, and determine the basic word range text1 - text2. Generate the Prompt for the input large model: You are a professional video copywriting generator. Please rewrite the current video {candidate materials} for the {current copywriting sentence being processed} into a video description within the word range of {text1 - text2} based on the theme of {the current copywriting sentence being processed}, and closely combine the context: the above text {the above text adjacent to the current copywriting sentence in the video copywriting}, the below text {the below text adjacent to the current copywriting sentence in the video copywriting}, and try to ensure that the copywriting does not deviate from the theme and the context logic is smooth. Use the large model to generate the latest copywriting within the word range text1 - text2.
[0153] For each copywriting sentence, construct a timeline according to the above operations, and then perform video rendering, adding TTS, BGM (background music), etc. to generate a patchwork video.
[0154] Among them, the construction method of the training data for the large model is as follows: for high-quality online video-text pairs, after splitting the copywriting into sentences, randomly select sentences as GT, and use GPT to generalize and rewrite the sentences into texts of different word counts as {the sentence of the copywriting currently being processed}. The sentences before and after are respectively used as {the previous sentence adjacent to the sentence of the copywriting currently being processed in the video copywriting} and {the next sentence adjacent to the sentence of the copywriting currently being processed in the video copywriting}. The video shot corresponding to the current sentence is used as {the candidate material} to generate the Prompt input to the large model. Calculate the text correlation score s1 between the output of the large model and GT, the consistency score s2 of adding the output of the large model to the context before and after the sentence of the copywriting currently being processed in the video copywriting, and calculate the correlation score s3 between the video shot corresponding to the current sentence and the output of the large model using gpt4o. Weighted sum the above three scores to obtain the loss function value, and train the large model based on the loss function value.
[0155] Based on the same principle as the Figure 1 method shown in Figure 7 FIG. shows a schematic structural diagram of a video generation device provided by an embodiment of the present disclosure, as Figure 7 shown, the video generation device 70 may include:
[0156] A preprocessing module 710, configured to obtain the text information to be processed, split the text information to be processed, and obtain a plurality of sentences to be processed;
[0157] A material acquisition module 720, configured to obtain at least one candidate material matching the sentence to be processed based on the sentence to be processed, and synthesize the voice to be processed corresponding to the sentence to be processed;
[0158] A material optimization module 730, configured to optimize the sentence to be processed when the video picture duration is less than the video voice duration, and generate a target video based on the optimized sentence to be processed and the candidate material;
[0159] Wherein, the video picture duration is the sum of the time lengths corresponding to all candidate materials, and the video voice duration is the time length corresponding to the voice to be processed.
[0160] In the video generation device provided by the embodiment of the present disclosure, when the video picture duration is less than the video voice duration, that is, when the candidate materials corresponding to the sentence to be processed are insufficient, the video picture and the video voice are made to correspond by optimizing the sentence to be processed, avoiding the worse the correlation between the video picture and the video voice or the video copywriting as the video timeline goes further back, improving the overall generation effect of the video, and avoiding video generation failure.
[0161] In some possible implementation manners, the material optimization module includes: a statistics sub-module, configured to count the number of reference words corresponding to the speech with the same duration as the video picture; an optimization sub-module, configured to optimize the sentence to be processed, so that the difference between the number of words included in the optimized sentence to be processed and the number of reference words is within a preset range; a speech sub-module, configured to generate the video speech corresponding to the optimized sentence to be processed based on the optimized sentence to be processed, compose the video picture based on the candidate materials, and generate the target video based on the video speech and the video picture.
[0162] In some possible implementation manners, the optimization sub-module includes: a word count unit, configured to determine the optimized word count range based on the number of reference words, where the number of reference words is within the optimized word count range, and the differences between the upper and lower bounds of the word count range and the number of reference words are both within the preset range; a large model unit, configured to construct a prompt text at least using the sentence to be processed and the optimized word count range, and input the prompt text into a pre-trained large model to obtain the optimized sentence to be processed with the number of words included within the optimized word count range.
[0163] In some possible implementation manners, the video generation device further includes a large model training module; the large model training module includes: a first training unit, configured to obtain a plurality of training videos and the training texts corresponding to the training videos, segment the training texts to obtain a plurality of true value sentences; a second training unit, configured to rewrite the true value sentences into training sentences with different numbers of words from the true value sentences; a third training unit, configured to construct a training prompt text at least using the training sentences and the training word count range, input the training prompt text into the large model, and train the large model according to the relevance between the output of the large model and the true value sentences; wherein, the number of words included in the true value sentences is within the training word count range; the differences between the upper and lower bounds of the training word count range and the number of words included in the true value sentences are both within the preset range.
[0164] In some possible implementation manners, the large model unit is configured to: construct a prompt text using the sentence to be processed, the optimized word count range, the candidate materials, the context adjacent to the sentence to be processed in the text information to be processed, and the following text.
[0165] In some possible implementation manners, the third training unit includes: constructing a training prompt text using the training sentences, the training word count range, the training video pictures of the training videos corresponding to the true value sentences, the context adjacent to the true value sentences in the training texts, and the following text; inputting the training prompt text into the large model, and training the large model according to the relevance between the output of the large model and the true value sentences, the consistency between the output of the large model and the context adjacent to the true value sentences in the training texts, and the relevance between the output of the large model and the training video pictures.
[0166] In some possible implementations, the material acquisition module includes: a retrieval sub-module for extracting keywords from the sentence to be processed and obtaining multiple retrieval materials that match the sentence to be processed based on the obtained keywords; a relevance sub-module for using a multi-modal large model to calculate the first relevance score between the retrieval materials and the sentence to be processed; a copywriting generation sub-module for using the multi-modal large model to generate copywriting based on the retrieval materials, obtaining the video copywriting corresponding to the retrieval materials, and calculating the second relevance score between the video copywriting and the sentence to be processed; a material screening sub-module for performing a weighted sum of the first relevance score and the second relevance score, and determining the retrieval materials as candidate materials when the weighted sum result is greater than a preset threshold.
[0167] In some possible implementations, the retrieval sub-module is used to: extract keywords from the sentence to be processed; obtain at least one video frame material that matches the sentence to be processed in the video material library based on the obtained keywords as the retrieval materials; obtain at least one picture material that matches the sentence to be processed in the picture material library based on the obtained keywords as the retrieval materials.
[0168] In some possible implementations, the video generation device further includes: a video generation module for cropping the video frame materials in the candidate materials to make the duration of the cropped video frames equal to the duration of the video voice when the duration of the video frames is greater than the duration of the video voice and the difference between the duration of the video frames and the duration of the video voice is greater than a preset duration threshold; modifying the time length of the picture materials corresponding to the candidate materials to make the duration of the modified video frames equal to the duration of the video voice when the duration of the video frames is greater than the duration of the video voice and the difference between the duration of the video frames and the duration of the video voice is not greater than a preset duration threshold.
[0169] It can be understood that the above-mentioned modules of the video generation device in the embodiments of the present disclosure have the functions of implementing the corresponding steps of the video generation method in the embodiments shown in Figure 1 The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated by multiple modules. For the function descriptions of the above modules of the video generation device, reference can be specifically made to the corresponding descriptions of the video generation method in the embodiments shown in Figure 1 The corresponding descriptions in the embodiments shown in are not repeated here.
[0170] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0171] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0172] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0173] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video generation method provided by the embodiment of the present disclosure.
[0174] Compared with the prior art, when the duration of the video picture is less than the duration of the video voice, that is, when the candidate materials corresponding to the sentence to be processed are insufficient, the sentence to be processed is optimized to make the video picture and the video voice correspond, avoiding that the later the video timeline is, the worse the correlation between the video picture and the video voice or the video copywriting is, improving the overall video generation effect, and avoiding video generation failure.
[0175] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the video generation method provided by the embodiment of the present disclosure.
[0176] Compared with the prior art, when the duration of the video picture is less than the duration of the video voice, that is, when the candidate materials corresponding to the sentence to be processed are insufficient, the sentence to be processed is optimized to make the video picture and the video voice correspond, avoiding that the later the video timeline is, the worse the correlation between the video picture and the video voice or the video copywriting is, improving the overall video generation effect, and avoiding video generation failure.
[0177] The computer program product includes a computer program, and the computer program, when executed by a processor, implements the video generation method provided by the embodiment of the present disclosure.
[0178] Compared with the prior art, when the duration of the video picture is less than the duration of the video voice, that is, when the candidate materials corresponding to the sentence to be processed are insufficient, the sentence to be processed is optimized to make the video picture and the video voice correspond, avoiding that the later the video timeline is, the worse the correlation between the video picture and the video voice or the video copywriting is, improving the overall video generation effect, and avoiding video generation failure.
[0179] Figure 8FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0180] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0181] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0182] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the video generation method in any other suitable manner (e.g., by means of firmware).
[0183] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0186] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0187] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0188] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0189] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0190] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video generation method, comprising: Acquire text information to be processed, and segment the text information to be processed to obtain a plurality of sentences to be processed; Based on the sentence to be processed, obtaining at least one candidate material matching the sentence to be processed, and synthesizing the speech to be processed corresponding to the sentence to be processed; When the duration of the video screen is less than the duration of the video voice, the sentence to be processed is optimized, and a target video is generated based on the optimized sentence to be processed and the candidate material; The video screen duration is the sum of the time lengths corresponding to all the candidate materials, and the video voice duration is the time length corresponding to the voice to be processed.
2. The method according to claim 1, wherein: The step of optimizing the sentence to be processed and generating a target video based on the optimized sentence to be processed and the candidate material includes: Counting the number of reference texts corresponding to the speech with the same duration as the video picture; Optimizing the sentence to be processed so that the difference between the number of characters included in the optimized sentence to be processed and the number of reference characters is within a preset range; Based on the optimized sentence to be processed, a video voice corresponding to the optimized sentence to be processed is generated, a video screen is composed based on the candidate materials, and a target video is generated based on the video voice and the video screen.
3. The method according to claim 2, wherein: The step of optimizing the sentence to be processed so that the difference between the number of characters included in the optimized sentence to be processed and the number of reference characters is within a preset range includes: Determine an optimized character quantity range based on the reference character quantity, the reference character quantity is within the optimized character quantity range, and the difference between the upper and lower bounds of the character quantity range and the reference character quantity is within a preset range; At least the sentence to be processed and the optimized character quantity range are used to construct a prompt text, and the prompt text is input into a pre-trained large model to obtain an optimized sentence to be processed whose number of characters is within the optimized character quantity range.
4. The method according to claim 3, further comprising: Acquire multiple training videos and training texts corresponding to the training videos, segment the training texts, and acquire multiple truth value sentences; Rewriting the truth-value sentence into a training sentence having a different number of characters from the truth-value sentence; At least construct a training prompt text using the training sentences and the training word quantity range, input the training prompt text into the large model, and train the large model according to the correlation between the output of the large model and the true value sentence; The number of characters included in the truth sentence is within the range of the training character number; the difference between the upper and lower bounds of the training character number range and the number of characters included in the truth sentence is within a preset range.
5. The method according to claim 4, wherein: The step of constructing a prompt text by at least using the sentence to be processed and the optimized word quantity range includes: The prompt text is constructed by using the sentence to be processed, the optimized word quantity range, the candidate materials, and the preceding and following texts adjacent to the sentence to be processed in the text information to be processed.
6. The method according to claim 5, wherein: The step of constructing a training prompt text by at least using the training sentences and the training word quantity range, inputting the training prompt text into the large model, and training the large model according to the correlation between the output of the large model and the true value sentence includes: Constructing a training prompt text using the training sentence, the training word quantity range, the training video screen of the training video corresponding to the truth sentence, and the preceding and following texts adjacent to the truth sentence in the training text; The training prompt text is input into the big model, and the big model is trained according to the relevance of the output of the big model with the true value sentence, the consistency of the output of the big model with the preceding and following contexts adjacent to the true value sentence in the training text, and the relevance of the output of the big model with the training video screen.
7. The method according to claim 1, wherein: The acquiring, based on the sentence to be processed, at least one candidate material matching the sentence to be processed comprises: Extracting keywords from the sentence to be processed, and acquiring multiple search materials matching the sentence to be processed based on the acquired keywords; Using a multimodal large model, calculating a first relevance score between the search material and the sentence to be processed; Using a multimodal large model, generating a text based on the search material, obtaining a video text corresponding to the search material, and calculating a second relevance score between the video text and the sentence to be processed; A weighted sum is performed on the first correlation score and the second correlation score, and when the weighted sum result is greater than a preset threshold, the search material is determined to be the candidate material.
8. The method according to claim 7, wherein: The step of extracting keywords from the sentence to be processed and obtaining a plurality of search materials matching the sentence to be processed based on the obtained keywords includes: Extracting keywords from the sentence to be processed; Based on the obtained keywords, at least one video screen material matching the sentence to be processed is obtained from a video material library as the search material; Based on the obtained keywords, at least one picture material matching the sentence to be processed is obtained from a picture material library as the search material.
9. The method according to claim 8, further comprising: When the video screen duration is greater than the video voice duration, and the difference between the video screen duration and the video voice duration is greater than a preset duration threshold, the video screen material in the candidate material is cropped so that the cropped video screen duration is equal to the video voice duration; When the video screen duration is greater than the video voice duration, and the difference between the video screen duration and the video voice duration is not greater than a preset duration threshold, modify the time length corresponding to the picture material in the candidate material so that the modified video screen duration is equal to the video voice duration.
10. A video generating device, comprising: A preprocessing module is used to obtain text information to be processed, and segment the text information to be processed to obtain a plurality of sentences to be processed; A material acquisition module, configured to acquire, based on the sentence to be processed, at least one candidate material matching the sentence to be processed, and synthesize the speech to be processed corresponding to the sentence to be processed; A material optimization module, used for optimizing the sentence to be processed when the duration of the video screen is shorter than the duration of the video voice, and generating a target video based on the optimized sentence to be processed and the candidate material; The video screen duration is the sum of the time lengths corresponding to all the candidate materials, and the video voice duration is the time length corresponding to the voice to be processed.
11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.