Video automatic generation method and device, equipment and medium

Through intelligent matching and recommendation technology combined with material library, video materials are generated and arranged, and the existing tools are solved, and the problems of inaccurate material matching, limited copywriting generation ability, high operation complexity, and insufficient personalization are achieved, and efficient and personalized video generation effects are achieved.

CN120050463APending Publication Date: 2025-05-27XIAMEN LIMAYAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510190396.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing automated video generation tools have problems such as inaccurate material matching, limited copywriting generation capabilities, high operational complexity, and insufficient personalization, which is difficult to meet the needs of ordinary creators to efficiently generate short videos.

Method used

By obtaining video-related information entered by the user, combining the preset material library, intelligent matching and recommendation technology are used to generate copywriting content, audio materials, visual materials and packaging materials that match user needs, and through the fusion and choreography mechanism of mixed and generated materials, the packaging elements are automatically integrated and applied to generate the final video.

Benefits of technology

It realizes accurate matching and intelligent recommendation of video materials according to user needs, significantly reducing video production time, improving production efficiency, and enhancing the accuracy, popularity and continuity of storyboard videos through personalized content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050463A_ABST
    Figure CN120050463A_ABST
Patent Text Reader

Abstract

The invention provides an automatic video generation method and device, equipment and a medium, and relates to the technical field of video production. The method comprises the following steps: acquiring video related information input by a user, and generating copywriting content according to the video related information and a preset copywriting script material library; according to a preset sound material library, converting the copywriting content into audio by using a text-to-voice conversion technology to obtain an audio material; according to the audio material, a digital human-driven technology is utilized, and a preset image material library and a video sub-mirror material library are combined, so that a visual material synchronized with the copywriting content is obtained; according to a preset packaging element material library, obtaining a packaging material matched with the copywriting content; and performing video synthesis and packaging on the audio material, the visual material and the packaging material to obtain a final video. According to the method, related materials can be generated through accurate matching and intelligent recommendation according to user requirements and recent popularity, a mixed cutting type and generation type material fusion arrangement mechanism is combined, one-key film formation is achieved, and the video production efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video production, and in particular, to a method, device, equipment and medium for automatically generating videos. Background Art

[0002] With the rise of short video platforms, more and more creators have started to pour into this industry. The production of short videos usually includes processes such as material collection, copywriting script writing, audio and video recording, and post-editing. This process takes a long time and requires high creative ability and professional knowledge of users. Therefore, there is an urgent need for a tool that can generate short videos with one click to meet the needs of ordinary creators to generate short videos efficiently.

[0003] Although there are some automated video generation tools in the existing market, they have the following deficiencies: (1) inaccurate material matching: lacking an intelligent recommendation mechanism and unable to accurately match relevant materials according to user needs. (2) Limited copywriting generation ability: Copywriting generation depends on preset templates, lacks creativity and flexibility, and it is difficult to generate high-quality and attractive content. (3) High operation complexity: Users need to manually handle multiple links, and the operation steps are cumbersome, affecting the production efficiency. (4) Lack of personalization: It is difficult to deeply customize according to the personal materials uploaded by users, and the video content lacks personalization and uniqueness.

[0004] In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention

[0005] The present invention aims to provide a method, device, equipment and medium for automatically generating videos to solve the shortcomings of inaccurate material matching, limited copywriting generation ability, high operation complexity, lack of personalization, etc. in the existing methods.

[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0007] A method for automatically generating videos, comprising:

[0008] S1, obtaining video-related information input by a user;

[0009] S2, generating copywriting content according to the video-related information and a preset copywriting script material library;

[0010] S3, according to a preset sound material library, using text-to-speech technology TTS to convert the copywriting content into audio to obtain audio materials;

[0011] S4, according to the audio materials, using digital human driving technology, combining a preset image material library and a video storyboard material library to obtain visual materials synchronized with the copywriting content;

[0012] S5. Obtain packaging materials that match the copy content according to a preset packaging element material library;

[0013] S6. Perform video synthesis and packaging on the audio material, the visual material, and the packaging material to obtain a finally synthesized and packaged video.

[0014] Preferably, update the copywriting script material library, the sound material library, the image material library, the video storyboard material library, and the packaging element material library in real time according to the video-related information; among them,

[0015] The copywriting script material library includes various copywriting templates, storyboard structures, and attractive catchphrases obtained by splitting the collected copywriting structures, storyboard structures, and catchphrases;

[0016] The sound material library includes common timbres and popular timbres;

[0017] The image material library includes pictures and emoticons in multiple fields;

[0018] The video storyboard material library includes storyboard video clips of different scenes;

[0019] The packaging element material library includes subtitles, special effects, and stickers.

[0020] Preferably, the specific content of S4 is as follows:

[0021] According to the copy content, combine the preset image material library and video storyboard material library, perform semantic similarity matching on each video shot and the copy content, and select the shots with the top r1 in semantic similarity ranking from the image material library and the video storyboard material library to obtain the first recalled shots;

[0022] From the video storyboard material library, obtain videos with a popularity higher than a set threshold in the recent n days, and combine the copy content to find popular videos with the top r2 in semantic similarity ranking;

[0023] Decompose and understand the popular video to obtain the shooting script of the popular video;

[0024] According to the shooting script, combine the preset image material library and video storyboard material library, perform matching and recommendation of shot materials with the same category element tags, and select the shots with the top r3 in element tag similarity ranking to obtain the second recalled shots;

[0025] According to the first recalled shots and the second recalled shots, obtain recommended shot materials;

[0026] Based on the copy content, the video storyboard materials are intelligently allocated according to the similarity between the understanding of the recommended video clip materials and each paragraph of the copy content, so as to obtain visual materials synchronized with the copy content.

[0027] Preferably, it further includes: when the similarity between the understanding of the recommended video clip materials and the current paragraph text of the copy content is lower than a preset threshold, the frame extraction map of the previously allocated video storyboard of the current paragraph text and its corresponding text content and the storyboard understanding are used as the current storyboard copy content, and a storyboard video synchronized with the current storyboard copy content is obtained as the current storyboard video.

[0028] Preferably, the video-related information includes: video application scenarios, context information, and user materials; when obtaining the audio material, the visual material, and the packaging material, if the user material contains content related to the corresponding material, and the content related to the material includes the audio uploaded by the user, the pictures and videos uploaded by the user, and the background stickers customized by the user, then the content related to the material in the user material is combined to optimize the audio material, the visual material, and the packaging material.

[0029] Preferably, based on the audio material, the visual material and the packaging material are respectively arranged and synthesized according to the semantic similarity with the audio material to obtain a preliminary video;

[0030] According to the packaging material, subtitles, special effects, and sticker effects are applied according to the semantic matching similarity with the storyboard understanding of the preliminary video to obtain the finally packaged video.

[0031] The present invention also provides a video automatic generation device, including:

[0032] An input unit for obtaining video-related information input by a user;

[0033] A copy content generation unit for generating copy content that meets requirements according to the video-related information and a preset copy script material library;

[0034] An audio material generation unit for converting the copy content into audio by using text-to-speech technology TTS according to a preset sound material library to obtain audio material;

[0035] A visual material generation unit for obtaining visual materials synchronized with the copy content by using digital human driving technology in combination with a preset image material library and a video storyboard material library according to the audio material;

[0036] A packaging material generation unit for obtaining packaging materials that match the copy content according to a preset packaging element material library;

[0037] A final video generation unit for performing video synthesis and packaging on the audio material, the visual material, and the packaging material to obtain a final video that is synthesized and packaged.

[0038] The present invention also provides a video automatic generation device, including a processor and a memory. A computer program is stored in the memory and can be executed by the processor to implement a video automatic generation method as described above.

[0039] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor of the device where the computer-readable storage medium is located, a video automatic generation method as described above is implemented.

[0040] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0041] By obtaining user input information and combining it with a preset material library, the present invention intelligently matches and recommends the generation of copy content, audio material, visual material, and packaging material related to the user input information. It combines the similarity to perform a mixed editing (i.e., intelligently allocates materials according to the similarity of the content of the two) and generative (if the similarity of the two is too low, a new shot is generated according to the previous shot) fusion arrangement for each material, automatically integrates and applies packaging elements, and generates a final video.

[0042] The present invention can accurately match and intelligently recommend the generation of relevant materials according to user needs and recent popularity, and can generate a complete video in one click, significantly reducing the video production time and improving the production efficiency. The present invention combines user personalized content, arranges video shots through the fusion of mixed editing and generative materials, and generates a video with high accuracy, popularity, and continuity between shot videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a schematic diagram of a video automatic generation method provided for Embodiment 1.

[0045] Figure 2 It is a schematic diagram of a video automatic generation device provided for Embodiment 2.

[0046] The present invention will be further described in detail below with reference to the drawings and specific embodiments. Specific Embodiments

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0048] Embodiment 1

[0049] Embodiment 1 of the present invention provides a method for automatically generating a video, which can be implemented by a video automatic generation device (hereinafter referred to as the generation device), and in particular, is executed by one or more processors in the generation device.

[0050] In this embodiment, the generation device may be an electronic device equipped with a processor, and the processor has a computer program for this video automatic generation method and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0051] As Figure 1 shown, a method for automatically generating a video includes steps S1 to S6.

[0052] S1. Obtain video-related information input by the user;

[0053] S2. Generate text content that meets the requirements according to the video-related information and a preset copywriting script material library;

[0054] S3. According to a preset voice material library, use text-to-speech technology TTS to convert the text content into audio to obtain audio materials;

[0055] S4. According to the audio materials, use digital human driving technology, combine a preset image material library and a video storyboard material library to obtain visual materials synchronized with the text content;

[0056] S5. According to a preset packaging element material library, obtain packaging materials that match the text content;

[0057] S6. Synthesize and package the audio material, the visual material, and the packaging material into a video to obtain the final video that has been synthesized and packaged.

[0058] In this embodiment, the user first selects a video production scenario, inputs context-related information, and uploads some materials. The system automatically generates a copywriting based on the video production scenario, the context-related information, and a preset copywriting script material library, and combines it with a preset sound material library to convert the copywriting into audio. At the same time, the system intelligently recommends and generates relevant images, video storyboards, and packaging elements from each material library. Through the mixed editing and generative material fusion and arrangement mechanism, all materials are intelligently arranged and then automatically synthesized and packaged into the final video clip, achieving the goal of generating high-quality videos with one click.

[0059] This embodiment mainly includes six modules:

[0060] (1) Input module:

[0061] Video application scenario selection: The user selects a corresponding scenario from the preset industry classifications (such as e-commerce, big health, emotion, law, automotive, etc.) and further selects the purpose of video production. Different video production purposes are set under each industry, corresponding to different video generation logics.

[0062] User context input information: According to the selected scenario, the user fills in key information such as product links, prices, selling points, pain points, usage scenarios, etc., and the system automatically associates relevant content.

[0063] User material upload: The user can upload personal materials, such as promotional pictures, display videos, copywriting, or audio files, to achieve personalized customization of the generated video.

[0064] In this embodiment, the video-related information includes: video application scenario, context information, and user materials.

[0065] (2) Material library module:

[0066] Copywriting script material library: Stores various copywriting templates, storyboard structures, and attractive catchphrases, and is dynamically updated to match hot topics. The material sources include copywriting and story structures, split shot architectures (storyboard design), and split catchphrases.

[0067] Sound material library: Stores various tone color materials, supports cloning of popular tone colors and dynamic updates. The sound material sources include common tone colors and cloning of popular tone colors.

[0068] Image material library: Manages pictures in multiple fields, emoticons, and user-uploaded materials, and continuously expands and updates. The material sources include commonly used emoticons in short videos, classic scene pictures, text-to-image (i.e., generating images from copywriting), image-to-image (i.e., generating images from images), etc.

[0069] Video Storyboard Footage Library: Stores storyboard video clips of different scenes, including those uploaded by users and automatically split. The sources of the footage include finding videos by scene, splitting videos, splitting videos from the user's personal asset library, producing public commercial video clips, and splitting existing footage through generative methods.

[0070] Packaging Element Footage Library: Contains elements required for video packaging such as subtitles, special effects, and stickers, and is updated regularly to conform to the popular trends. The sources of the footage include copyrighted fonts, music, special effects, etc.

[0071] (3) Footage Processing Module:

[0072] Copywriting Processing: Analyzes the structure of the copy input by the user, extracts the key parts, and optimizes the content in combination with the preset copywriting script footage library. It supports ways such as copying the structure, beginning, and riding on hot topics of existing copywriting script footage.

[0073] Sound Processing: Performs voice cloning to generate natural and smooth narration audio.

[0074] Image Processing: Based on the existing image footage library and the picture footage uploaded by the user, it intelligently recommends and generates relevant pictures and emoticons to enhance the visual effect. For example, it can recommend pictures based on similarity; it can generate relevant pictures with the help of existing large models.

[0075] (4) Footage Generation Module:

[0076] Copywriting Content Generation: Generates copywriting content that meets the requirements based on the context information and the copywriting script footage library.

[0077] Audio Footage Generation: Utilizes TTS technology (Text-to-Speech), which is text-to-speech conversion technology, to convert the generated copywriting content into audio footage.

[0078] Visual Footage Generation: Based on the audio footage, combined with the preset image footage library and video storyboard footage library, it generates or matches visual footage that is synchronized with the copywriting content through digital human driving technology. In this embodiment, the digital human driving technology refers to the technology that drives digital humans to perform various activities and interactions. It combines multiple technologies such as artificial intelligence, computer graphics, motion capture, and speech synthesis, and is mainly divided into two categories: intelligent driving and real-person driving.

[0079] Intelligent Driving: The intelligent system automatically reads and parses the externally input information (such as text), makes decisions on the subsequent output text of the digital human according to the parsing results, and then drives the character model (referred to as the TTSA character model) pre-trained through AI technology to generate corresponding voices and actions to interact with the user.

[0080] Human-driven: The expressions and movements of a real person are collected through a motion capture system and presented on the virtual digital human image. At the same time, the real person communicates with the user in real time through voice, thus realizing interaction with the user.

[0081] If the user materials contain content related to the corresponding materials, that is, including the audio uploaded by the user, the pictures and videos uploaded by the user, and the background map customized by the user, then the content related to the materials in the user materials is used as one of the materials to optimize the generated audio to meet the requirements of personalized customization.

[0082] (5) Material recommendation and arrangement module:

[0083] Hybrid editing and generative material fusion and arrangement mechanism: Combining pre-existing materials (the copywriting script material library, the sound material library, the image material library, the video storyboard material library, and the packaging element material library) and generated materials (copywriting content, audio materials, visual materials, and the packaging materials), intelligently recommend and arrange various materials, including timbre, pictures, video storyboard segments, etc. Adopting a combination of hybrid editing and generative methods to achieve the diversity and continuity of materials.

[0084] For example, when generating visual materials, according to the copywriting content, combining the preset image material library, the video storyboard material library, and the user materials, perform semantic similarity matching on each video shot and the copywriting content, and select the shots with the top r1 (such as the top 10) semantic similarity rankings from the image material library, the video storyboard material library, and the user materials to obtain the first recalled shots.

[0085] From the video storyboard material library, obtain videos with a popularity higher than a set threshold (such as 90%) in the recent n days (such as the recent 7 days), and combine the copywriting content to find the popular videos with the top r2 (such as the top 3) semantic similarity rankings.

[0086] Decompose and understand the popular video to obtain the shooting script of the popular video. This step is to make the generated video have the elements of the popular video.

[0087] According to the shooting script, combining the preset image material library, the video storyboard material library, and the user materials, perform matching and recommendation of shot materials with the same category element tags, and select the shots with the top r3 (such as the top 5) element tag similarity rankings to obtain the second recalled shots.

[0088] According to the first recalled shots and the second recalled shots, obtain the recommended shot materials.

[0089] In this embodiment, the values of r1, r2, r3, and n can be set according to the actual needs of the user and are not limited here.

[0090] Based on the content of the copywriting, the video storyboard materials are intelligently allocated according to the similarity between the understanding of the storyboard of the recommended video materials and each paragraph of the copywriting content, and visual materials synchronized with the copywriting content (i.e., mixed editing style) are obtained.

[0091] When the similarity between the understanding of the storyboard of the recommended video materials and the current paragraph of the copywriting content is lower than the preset threshold, the frame extraction map of the previously allocated video storyboard of the current paragraph of the copywriting and its corresponding text content and the storyboard understanding are used as the current storyboard copywriting content, and a storyboard video synchronized with the current storyboard copywriting content is generated as the current storyboard video (i.e., generative style).

[0092] Material arrangement: According to the audio duration and content, the recommended materials are arranged into the video by the above-mentioned mixed editing and generative methods to ensure the fluency and attractiveness of the video content.

[0093] (6) Video synthesis and packaging module:

[0094] Video synthesis: Integrate sound, images, video storyboards and packaging elements to generate a preliminary video. Specifically, based on the audio materials, the visual materials and the packaging materials are respectively arranged and synthesized according to the semantic similarity with the audio materials to obtain a preliminary video.

[0095] Video packaging: Apply packaging elements such as subtitles, special effects, and stickers to enhance the visual effect and professionalism of the video.

[0096] Output the final video: Generate a final video file containing all packaging elements, and users can directly download or share it.

[0097] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0098] (1) Overall process design: The present invention covers a fully automated process from material collection, processing, generation to video synthesis, optimizes the connection and cooperation of each link, and improves the efficiency and quality of video production.

[0099] (2) Design of each material library and material sources: The present invention contains a systematic design of multiple material libraries, including copywriting scripts, sounds, images, video storyboards, and packaging elements, etc., to ensure the richness and diversity of materials. The material sources cover copyrighted materials, user-uploaded materials, and fission by generative methods, ensuring the legality and innovation of materials. At the same time, it supports users to upload personal materials and generates highly personalized video content according to user needs, enhancing the user experience.

[0100] (3) Mixed - editing and generative material fusion and arrangement mechanism: The present invention combines the mixed - editing and generative methods to achieve intelligent recommendation and matching arrangement of materials, ensuring the continuity and diversity of video content. By using intelligent algorithms, materials of different types such as text content matching, audio, vision, and packaging elements are fused, improving the accuracy of material matching, the continuity and attractiveness of video content, and enhancing the overall effect of the video.

[0101] (4) Continuous update: The present invention supports dynamic update of the material library to ensure that the system content keeps up with the times and meets the diverse and changing market demands.

[0102] Embodiment Two

[0103] As Figure 2 shown, the second embodiment of the present invention also provides a video automatic generation device, including:

[0104] An input unit, configured to obtain video - related information input by a user;

[0105] A text content generation unit, configured to generate text content that meets the requirements according to the video - related information and a preset text script material library;

[0106] An audio material generation unit, configured to convert the text content into audio using text - to - speech technology TTS according to a preset sound material library to obtain audio materials;

[0107] A visual material generation unit, configured to generate visual materials synchronized with the text content according to the audio materials, using digital human driving technology and combining a preset image material library and a video storyboard material library;

[0108] A packaging material generation unit, configured to obtain packaging materials that match the text content according to a preset packaging element material library;

[0109] A final video generation unit, configured to perform video synthesis and packaging on the audio materials, the visual materials, and the packaging materials to obtain a synthesized and packaged final video.

[0110] Embodiment Three

[0111] The third embodiment of the present invention also provides a video automatic generation device, which includes a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement the video automatic generation method as described above.

[0112] Embodiment Four

[0113] The fourth embodiment of the present invention further provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the above-described video automatic generation method is implemented.

[0114] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0115] In addition, the various functional modules in the embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0116] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-On l y Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0117] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0118] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0119] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0120] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged with a specific order or sequence under allowable circumstances. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for automatically generating a video, characterized in that: include: S1, obtaining video related information input by the user; S2, generating copy content according to the video related information and a preset copy script material library; S3, according to a preset sound material library, using the text-to-speech conversion technology TTS, converting the text content into audio to obtain audio material; S4, based on the audio material, using the digital human driving technology, combined with a preset image material library and a video storyboard material library, to obtain visual materials synchronized with the content of the copy; S5, obtaining packaging materials matching the content of the copy according to a preset packaging element material library; S6, synthesizing and packaging the audio material, the visual material and the packaging material to obtain a synthesized and packaged final video.

2. A method for automatically generating a video according to claim 1, characterized in that , update the copywriting script material library, the sound material library, the image material library, the video storyboard material library and the packaging element material library in real time according to the video related information; wherein, The copywriting script material library obtains various copywriting templates, storyboard structures and attractive golden sentences by splitting and collecting copywriting structures, storyboard structures and golden sentences; The sound material library includes commonly used timbres and popular timbres; The image material library includes pictures and emoticons from multiple fields; The video storyboard material library includes storyboard video clips of different scenes; The packaging element library includes subtitles, special effects and textures.

3. A method for automatically generating a video according to claim 1, characterized in that , the S4 is specifically: According to the content of the copy, combined with the preset image material library and video storyboard material library, each video shot is matched with the content of the copy for semantic similarity, and the shots with the top r1 semantic similarity ranking are selected from the image material library and the video storyboard material library to obtain the first recalled shot; From the video storyboard library, obtain videos whose popularity in the past n days is higher than a set threshold, and find the top r2 popular videos with the highest semantic similarity in combination with the copy content; Decomposing and understanding the hit video to obtain a shooting script of the hit video; According to the shooting script, combined with the preset image material library and video storyboard material library, the lens materials with the same category of element labels are matched and recommended, and the lenses with the top r3 element label similarity rankings are selected to obtain the second recalled lenses; Obtaining recommended shot materials according to the first recalled shot and the second recalled shot; Based on the content of the copy, the video storyboard material is intelligently allocated according to the similarity between the storyboard understanding of the recommended shot material and each paragraph of the text content, so as to obtain the visual material synchronized with the content of the copy.

4. A method for automatically generating a video according to claim 3, characterized in that: Also includes: If the similarity between the storyboard understanding of the recommended shot material and the current paragraph text of the copy content is lower than a preset threshold, the frame extraction image of the previous assigned video storyboard of the current paragraph text and its corresponding text content and storyboard understanding are used as the current storyboard copy content to obtain a storyboard video synchronized with the current storyboard copy content as the current storyboard video.

5. A method for automatically generating a video according to claim 4, characterized in that , the video-related information includes: video application scenarios, context information and user materials; when obtaining the audio material, the visual material and the packaging material, if the user material contains content related to the corresponding material, the material-related content includes audio uploaded by the user, pictures and videos uploaded by the user, and background maps customized by the user, then the audio material, the visual material and the packaging material are optimized in combination with the material-related content in the user material.

6. The method for automatically generating a video according to claim 1, characterized in that: The S6 is specifically: Based on the audio material, the visual material and the packaging material are arranged and synthesized according to the semantic similarity with the audio material to obtain a preliminary video; Based on the packaging material, subtitles, special effects and texture effects are applied according to the semantic matching similarity with the storyboard understanding of the preliminary video to obtain a packaged final video.

7. A video automatic generation device, characterized in that: include: An input unit, used to obtain video related information input by a user; A copy content generating unit, used to generate copy content that meets the requirements according to the video related information and a preset copy script material library; An audio material generating unit, used to convert the text content into audio according to a preset sound material library by using the text-to-speech conversion technology TTS to obtain audio material; A visual material generation unit, configured to obtain visual materials synchronized with the content of the text based on the audio material, using digital human driving technology, and combining a preset image material library and a video storyboard material library; A packaging material generating unit, used to obtain packaging materials matching the content of the copy according to a preset packaging element material library; The final video generation unit is used to perform video synthesis and packaging on the audio material, the visual material and the packaging material to obtain a synthesized and packaged final video.

8. The video automatic generation device according to claim 7, characterized in that: The visual material generation unit is specifically: According to the content of the copy, combined with the preset image material library and video storyboard material library, each video shot is matched with the content of the copy for semantic similarity, and the shots with the top r1 semantic similarity ranking are selected from the image material library and the video storyboard material library to obtain the first recalled shot; From the video storyboard library, obtain videos whose popularity in the past n days is higher than a set threshold, and find the top r2 popular videos with the highest semantic similarity in combination with the copy content; Decomposing and understanding the hit video to obtain a shooting script of the hit video; According to the shooting script, combined with the preset image material library and video storyboard material library, the lens materials with the same category of element labels are matched and recommended, and the lenses with the top r3 element label similarity rankings are selected to obtain the second recalled lenses; Obtaining recommended shot materials according to the first recalled shot and the second recalled shot; Based on the content of the copy, the video storyboard material is intelligently allocated according to the similarity between the storyboard understanding of the recommended shot material and each paragraph of the text content, so as to obtain the visual material synchronized with the content of the copy.

9. A video automatic generation device, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a video automatic generation method as described in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a method for automatically generating a video as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Digital human material generation and sharing method and device based on intelligent hardware

    CN120640079A

  • Digital human material generation and sharing method and device based on intelligent hardware

    CN120640079B