Automatic content generation method and device, electronic equipment and storage medium

By using large language models and AI technology to generate scripts, images, and audio and then blending them in style, the problem of low content creation efficiency and poor style consistency in existing technologies has been solved. This has enabled efficient and diverse video content generation, improving the quality and adaptability of the content.

CN120897099APending Publication Date: 2025-11-04ZHAOLIAN CONSUMER FINANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510800070.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies suffer from low content creation efficiency, poor cross-modal content integration, and poor stylistic consistency, resulting in cumbersome creation processes and difficulty in quickly adapting to changes in market demands.

Method used

The script is generated through deep semantic analysis using a large language model. A pre-trained script generation model is used to generate video scripts. High-quality images are generated by combining AI technology and image generation algorithms. Text-to-speech technology optimizes the speech, and style fusion technology ensures consistency. Finally, video synthesis technology is used to edit and generate multiple versions of video content.

Benefits of technology

It enables automated, diversified, and high-quality video content generation, solving problems such as low generation efficiency, poor cross-modal content integration, and poor style consistency, thereby improving creation efficiency and quality and meeting the needs of different users and platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897099A_ABST
    Figure CN120897099A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an automatic content generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence content generation, the method comprises the following steps: obtaining a demand text, using a large language model to carry out deep semantic analysis on the demand text to obtain a script, and using a script generation model to generate a video script; performing semantic understanding through an A I technology and an image generation algorithm to generate an image, and converting a script into voice through a text-to-voice technology and a deep learning model; performing style analysis on the script, the image and the voice to extract style features, and performing style fusion according to the style features and style requirements in the demand text; and aligning the script, the image and the voice according to a time sequence, and editing and integrating the script, the image and the voice by using a video synthesis technology to generate a plurality of versions of video contents. According to the method, the problems of low generation efficiency, poor cross-modal content fusion and poor style uniformity in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence content generation, and in particular to an automatic content generation method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of Internet technology, the field of content creation has undergone unprecedented changes. Traditional content creation methods, such as manually writing scripts, manually generating images and video clips, are not only time-consuming and labor-intensive, but also difficult to meet the rapidly changing market demand. Although existing technologies have attempted to use AI to assist content creation, there are still problems such as low generation efficiency, poor cross-modal content coordination, and insufficient system scalability.

[0003] Traditional content generation technology relies on multiple scattered tools and platforms, resulting in a cumbersome creation process, long data flow links, and the need for multiple manual interventions. For example, script creation relies on manual writing or simple templates, and lacks AI-assisted creative optimization; image, audio and video generation are scattered in different systems, requiring manual switching and adjustment, which greatly reduces the creation efficiency.

[0004] In addition, the content style and consistency generated by different AI tools are difficult to guarantee, and a large amount of manual optimization is required in the later stage. Existing systems are mostly customized, making it difficult to quickly adapt to changes in new scenarios and requirements, and lacking modular design, which limits the scalability and reusability of the system.

[0005] Therefore, there is an urgent need for an automatic content generation method that can efficiently integrate all aspects of content generation, achieve content style uniformity and cross-modal fusion, and thus improve content creation efficiency and quality. SUMMARY

[0006] Embodiments of the present application provide an automatic content generation method to solve the problems of low generation efficiency, poor cross-modal content fusion and poor style uniformity in existing technologies. The technical solution is as follows:

[0007] According to one aspect of the present application, an automated content generation method, the method comprising: obtaining a requirement text, using a large language model to perform deep semantic analysis on the requirement text to obtain a script, using a pre-trained script generation model to generate a video script according to the script; generating images according to the script and the video script after semantic understanding by AI technology and image generation algorithm, converting the script into voice by text-to-speech technology and deep learning model; performing style analysis on the script, images and voice to extract style features, and performing style fusion on the script, images and voice according to the style features and style requirements in the requirement text; aligning the script, images and voice in chronological order, and editing and integrating the script, images and voice using video synthesis technology to generate multiple versions of video content.

[0008] In one embodiment, using a large language model to perform deep semantic analysis on the requirement text to obtain a script, and using a pre-trained script generation model to generate a video script is achieved by the following steps: using the large language model to extract key information, intent and emotional color in the requirement text to obtain semantic information, and using a pre-trained script generation model to generate a structured script according to the semantic information; the requirement text includes video style, target audience, emotional tone, background environment and character setting; performing semantic analysis and structured processing on the script to extract key elements to generate a video script; the key elements include scene, dialogue, rhythm, emotion, duration and shot change information; the video script includes scene description, character dialogue and action guidance.

[0009] In one embodiment, generating images according to the script and the video script after semantic understanding by AI technology and image generation algorithm is achieved by the following steps: performing semantic understanding on the script and the video script, extracting key visual elements and scene descriptions, and generating high-quality images according to the visual elements and scene descriptions by AI technology and image generation algorithm; the images include multiple style and resolution versions; the image generation algorithm includes a generative adversarial network.

[0010] In one embodiment, converting the script into voice by text-to-speech technology and deep learning model is achieved by the following steps: converting the text content in the script into natural and fluent voice by text-to-speech technology, optimizing the emotional color of the voice by deep learning model, and adjusting the timbre, speed and tone of the voice.

[0011] In one of the embodiments, the style analysis is performed on the script, image and voice to extract style features, and the style fusion is performed on the script, image and voice according to the style requirements in the requirement text by the following steps: analyzing the style of the script, image and voice to extract key style features, and performing style matching on most of the image, voice and text according to the style requirements in the requirement text; and performing fusion optimization on the image, voice and text so that the script, image and voice are consistent in style; the fusion optimization includes adjusting the color tone of the image, the tone of the voice and the layout of the text.

[0012] In one of the embodiments, the video synthesis technology is used to edit and integrate the script, image and voice to generate multiple versions of video by the following steps: editing the aligned image, voice and text, synthesizing the image, voice and text by using the video synthesis technology to obtain a video, previewing the video and exporting it in multiple formats; the editing includes adding transition effects, background music and subtitles.

[0013] In one of the embodiments, the method further includes the following steps: using AI technology to perform material matching, material replacement, adding transition and adding special effects for the video to generate multiple versions of video, and automatically detecting the generated video to delete the video with quality problems; the quality problems include audio and picture out of sync, black frame and frame skipping.

[0014] According to one aspect of the present application, an automatic content generation device, the device includes: a video script generation module for obtaining a requirement text, using a large language model to perform deep semantic analysis on the requirement text to obtain a script, and using a pre-trained script generation model to generate a video script according to the script; an image and voice conversion module for generating an image according to the script and the video script through AI technology and image generation algorithm after semantic understanding, and converting the script into voice through text-to-speech technology and a deep learning model; a style fusion and unification module for performing style analysis on the script, image and voice to extract style features, and performing style fusion on the script, image and voice according to the style features and the style requirements in the requirement text; and a video editing and synthesis module for aligning the script, image and voice in time sequence, and using video synthesis technology to edit and integrate the script, image and voice to generate multiple versions of video content.

[0015] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory has computer readable instructions stored thereon; the computer readable instructions are executed by one or more processors, so that the electronic device implements the automatic content generation method as described above.

[0016] According to an aspect of the present application, a storage medium having stored thereon computer readable instructions to be executed by one or more processors to implement the automated content generation method as described above.

[0017] The technical solution provided by the present application has the following beneficial effects:

[0018] In the above technical solution, the present application first acquires a demand text, generates a script by means of deep semantic analysis of a large language model, and then generates a video script according to the script by using a pre-trained script generation model. Next, semantic understanding is performed on the script and the video script by means of AI technology and an image generation algorithm to generate high-quality images containing multiple styles and resolutions. At the same time, the script is converted into natural and fluent speech with optimized emotion, timbre, speech speed, and intonation by using text-to-speech technology and a deep learning model. Then, style analysis is performed on the script, images, and speech, key style features are extracted, and style fusion and optimization are performed on the three according to the style requirements in the demand text to ensure style consistency. Subsequently, the script, images, and speech are aligned in chronological order, and video synthesis technology is used for editing and integration to generate multiple versions of videos. The editing process covers adding transition effects, background music, subtitles, etc. In addition, AI technology is used for material matching, replacement, and addition of transition and special effects to generate more versions of videos, and videos with quality problems such as audio-visual asynchronization, black frames, and frame skipping are automatically detected and deleted. The automated, diversified, and high-quality video content generation realizes automatic generation, diversification, and high quality, effectively solving the problems of low generation efficiency, poor cross-modal content fusion, and poor style uniformity in the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0020] Figure 1 is a flowchart of an automated content generation method according to an exemplary embodiment;

[0021] Figure 2 is a flowchart of an automated content generation method according to an exemplary embodiment;

[0022] Figure 3 is a flowchart of training a fawn model in an application scenario;

[0023] Figure 4 is a block diagram of an automated content generation device according to an exemplary embodiment;

[0024] Figure 5 FIG. 1 is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0025] Figure 6 FIG. 2 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0026] Embodiments of the present application are described in detail below with reference to the attached drawings, which show by way of example, embodiments in which like numerals indicate like elements or elements having the same or similar function throughout the several disclosed embodiments. The embodiments described below are exemplary only and are not to be construed as limiting the present application.

[0027] As can be appreciated by those skilled in the art, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further understood that the terms "comprise," "comprises," "comprising," "include," "includes," "including," "contain," "contains," "containing," and the like, when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements can be present. In addition, the use herein of "connected" or "coupled" can include wireless connection or wireless coupling. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0028] The present application provides an automatic content generation method, which automatically generates video content of multiple styles and high quality by acquiring a demand text and using large language models, AI technology, image generation algorithms, text-to-speech technology, and other means, while automatically detecting and deleting problem videos, effectively solving the problems of low efficiency, difficulty in unifying styles, and difficulty in quality control in traditional video generation. The automatic content generation method is suitable for use in an automatic content generation device, which can be an electronic device. The automatic content generation method in the embodiments of the present application can be applied to various scenarios, such as self-media video generation, etc.

[0029] Referring to Figure 1 The embodiments of the present application provide an automatic content generation method, which is suitable for use in an electronic device.

[0030] In the following method embodiments, for the convenience of description, the execution subject of each step of the method is taken as an example to be described as an electronic device, but this does not constitute a specific limitation.

[0031] As Figure 1As shown, the method can include the following steps:

[0032] In step 110, the demand text is obtained, a large language model is used to perform deep semantic analysis on the demand text to obtain a script, and a pre-trained script generation model is used to generate a video script according to the script.

[0033] In one possible implementation, the key information, intent and emotional color in the demand text are extracted using a large language model to obtain semantic information, a pre-trained script generation model is used to generate a structured script according to the semantic information, and the script is analyzed and structured to extract key elements to generate a video script.

[0034] Among them, the demand text includes video style, target audience, emotional tone, background environment and character setting, etc., the key elements include scene, dialogue, rhythm, emotion, duration and shot switching information, etc., and the video script includes scene description, character dialogue and action guidance, etc., which are not limited here.

[0035] In the above process, by obtaining comprehensive demand text, the system can accurately understand the user's creative intent to provide a solid foundation for subsequent steps, deep semantic analysis ensures that the script accurately reflects user demand, and the structured script provides a clear framework for video script generation. The generation of the video script makes the video creation process more orderly, improving the efficiency and quality of the creation.

[0036] In step 120, images are generated according to the script and the video script after semantic understanding by AI technology and image generation algorithms, and the script is converted to speech by text-to-speech technology and a deep learning model.

[0037] In one possible implementation, the script and the video script are semantically understood, key visual elements and scene descriptions are extracted, and high-quality images are generated according to the visual elements and scene descriptions by AI technology and image generation algorithms.

[0038] Among them, the images include various styles and resolution versions, etc., and the image generation algorithm includes a generative adversarial network, etc., which are not limited here.

[0039] In one possible implementation, the text content in the script is converted to natural and fluent speech by text-to-speech technology, the emotional color of the speech is optimized by a deep learning model, and the timbre, speed and tone of the speech are adjusted.

[0040] In the above process, the embodiment of the present application provides rich visual materials for video creation by generating high-quality images, meets different scenes and needs through various styles and resolutions, enhances the immersion and realism of the video through natural and smooth voice, and makes the voice more consistent with the video content through optimization of emotional color and adjustment of voice parameters.

[0041] In step 130, the script, images, and voice are analyzed for style to extract style features, and the script, images, and voice are style-fused according to the style requirements in the demand text.

[0042] In one possible implementation, the style of the script, images, and voice is analyzed to extract key style features, and the majority of images, voice, and text are style-matched according to the style requirements in the demand text, and the images, voice, and text are fused and optimized to make the script, images, and voice consistent in style.

[0043] The fusion and optimization includes adjusting the color tone of the images, the tone of the voice, and the layout of the text, etc., which are not limited here.

[0044] In the above process, the embodiment of the present application ensures the coordination and consistency of the video content in style through style matching and fusion optimization, and improves the overall quality of the video.

[0045] In step 140, the script, images, and voice are aligned in time sequence, and video synthesis technology is used to edit and integrate the script, images, and voice to generate multiple versions of video content.

[0046] In one possible implementation, the aligned images, voice, and text are edited, and video synthesis technology is used to synthesize the images, voice, and text to obtain a video and preview it, and export it in multiple formats.

[0047] In one possible implementation, AI technology is used to perform material matching, material replacement, add transitions, and add special effects to generate multiple versions of video, and automatically detect the generated video to delete the video with quality problems.

[0048] The editing includes adding transition effects, background music, subtitles, etc., and the quality problems include audio and picture out of sync, black frames, and frame skipping, etc., which are not limited here.

[0049] In the above process, the embodiment of the present application makes the video content more coherent and rich through material alignment and editing, improves the viewing experience of the audience, video synthesis and export makes the video content can be easily played and shared on different platforms, video optimization and quality detection ensures that the finally generated video content has high quality and stability, and improves the user's satisfaction.

[0050] Through the above process, the present application realizes the automatic generation from the demand text to the high-quality video content by demand analysis and script generation, multimedia material generation, style fusion processing, and video synthesis and optimization, has the characteristics of high efficiency, accuracy, flexibility and strong scalability, can meet the video creation needs of different users in different scenarios, the generated video content is consistent in style and stable and reliable in quality, and brings revolutionary changes to the video creation field.

[0051] In an exemplary embodiment, multiple versions of videos need to be quickly generated to adapt to different platforms and audience groups, demonstrating the process of automated video generation.

[0052] As Figure 2 shown, it can specifically include the following steps:

[0053] Step one: input the video script.

[0054] Specifically, the scriptwriter inputs the carefully written video script into the automated content generation system, and the script content covers information such as content highlights and scene descriptions. The input script serves as the starting point of the entire generation process and provides a basic direction for subsequent processing, ensuring that the generated video can accurately convey the core information.

[0055] Step two: semantic analysis and structuring of the script, and extraction of key elements.

[0056] Specifically, the system uses natural language processing technology to perform in-depth semantic analysis on the script, structures it, and extracts key elements such as key objects, scene types, and character actions. Through accurate semantic analysis, the core content of the script can be accurately grasped, and the extracted key elements provide accurate basis for subsequent material matching and video generation, improving the fit degree of the generated video and the script.

[0057] Step three: determine whether it is a multi-material library.

[0058] Specifically, the system determines whether the required material comes from multiple material libraries according to the preset rules. This judgment link can flexibly select the material acquisition method according to the actual situation. If it is a multi-material library, the rich material resources can be fully utilized to increase the diversity and richness of the video; if it is a single material library, the material calling process is simplified, and the generation efficiency is improved.

[0059] Step four: multi-material library intelligent matching or single-material library calling.

[0060] Specifically, if it is a multi-material library, the system uses intelligent algorithms to filter suitable materials from multiple material libraries based on the extracted key elements. Intelligent matching can consider the resources of multiple material libraries to find the most suitable materials for the script requirements, making the video more rich and high-quality in terms of picture and sound effects, and improving the appeal.

[0061] Specifically, if it is a single-material library, the system directly calls related materials from the material library. The single-material library calling process is simple, which can quickly obtain materials and improve the generation speed of the video, and is suitable for production scenarios with high time requirements.

[0062] Step five: AI material screening and dynamic matching, automatic editing.

[0063] Specifically, the system further uses AI technology to screen and dynamically match the selected materials, and automatically edits according to the structure and rhythm of the script. AI material screening can remove unsuitable materials, dynamic matching ensures smooth connection between materials, and automatic editing greatly shortens the video production time while ensuring the quality and coherence of the video.

[0064] Step six: Determine whether video fission is needed.

[0065] Specifically, the system determines whether multiple versions of video variants need to be generated based on preset rules or user needs. This judgment can meet the personalized needs of different platforms and audience groups. If video fission is needed, multiple styles of videos can be generated to expand coverage and influence. If not, a single version of the refined video can be directly output to improve production efficiency.

[0066] Step seven: Replace materials or transitions, batch generate video variants or single version refined output.

[0067] Specifically, if video fission is needed, the system generates multiple versions of video variants by replacing some materials or adjusting the transition method. Batch generation of video variants can quickly adapt to the characteristics of different platforms and audience preferences, improving the effectiveness of placement and conversion rate.

[0068] Specifically, if video fission is not needed, the system outputs a single version of the refined video. Single version refined output can ensure the high quality and professionalism of the video, while reducing unnecessary production costs and time waste.

[0069] Step eight: Automatically detect and output the generated video.

[0070] Specifically, the system automatically detects the generated video, checks whether there are problems such as audio and picture out of sync, black frames, frame skipping, and the like, and outputs the final video after confirmation. The automatic detection link ensures the stability of the video quality, avoids the influence of video quality problems on the effect, and the output of high-quality videos can better show the product image and propaganda information.

[0071] Through the above process, the embodiment of the application realizes the generation of automatic video. The entire process is started by inputting a video script, key elements are extracted by semantic analysis and structuring, appropriate material acquisition methods are selected according to the material library, and automatic editing is completed through AI material screening and dynamic matching. Whether to perform video fission is determined according to the demand, and finally the generated video is detected and output. The entire process fully utilizes the advantages of automatic technology, improves the production efficiency and quality of the video, quickly responds to market demand, generates diversified and high-quality videos, and brings significant benefit improvement and innovative development to the content production industry.

[0072] In an application scenario, a user wants to quickly and conveniently create a video with personalized style, and uses the automatic content generation method provided by the application to generate automatic content.

[0073] Step one: the user initiates creation.

[0074] Specifically, the user initiates a video creation request on the operation interface, which is the starting point of the entire video creation process. The simple and convenient initiation method allows the user to easily start creation, lowers the creation threshold, and enables more users to participate in video creation.

[0075] Step two: the content generation platform returns a script and style setting.

[0076] Specifically, after receiving the user's creation request, the content generation platform creates a script and returns it to the user. At the same time, the user can choose to set the style and quantity. In addition, the platform also provides functions such as expanding the script, abbreviating the script, and changing the script style. The content generation platform quickly generates a script, providing a basic framework for user creation.

[0077] Among them, the user's setting of style and quantity can meet the personalized needs, while the functions of expanding, abbreviating and changing the style of the script further enhance the flexibility of creation, making the script more in line with the user's expectations.

[0078] Step three: the user generates a video request.

[0079] Specifically, after confirming the script and style settings, the user issues a video generation instruction to the content generation platform. The explicit generation instruction triggers the subsequent video generation process, ensuring that the entire creation process proceeds in an orderly manner.

[0080] Step four: the content generation platform provides the script with a scene and returns the scene with the script.

[0081] Specifically, the content generation platform matches the script with a corresponding scene according to the script content and returns the script with scene information, which enriches the video content performance by matching the script with a scene, making the video more visually appealing and situational, and improving the video's watchability.

[0082] Step five: the content generation platform analyzes the scene to generate SD instructions and responds.

[0083] Specifically, the content generation platform analyzes the script with the scene to generate Stable Diffusion (SD) instructions and uses the Zilu large language model to respond to these instructions. By analyzing the scene to generate SD instructions, abstract scene descriptions can be converted into specific image generation instructions, providing precise guidance for subsequent image generation and ensuring that the generated images meet the script scene requirements.

[0084] Among them, the Zilu large model uses high-precision semantic recognition and analysis functions, combined with context, to refine and expand, generating video scripts containing rich plots and attractive dialogues. At the same time, according to the scene limitations customized by the operation side, excellent video scripts are quickly generated, and natural language processing technology is used to realize automatic analysis and understanding of video scripts, supporting automatic generation and optimization of scripts.

[0085] Step six: AI text-to-image generates pictures and returns them.

[0086] Specifically, the AI text-to-image system generates pictures based on the received SD instructions and returns the generated pictures to the content generation platform. Using AI text-to-image technology, high-quality pictures can be quickly and efficiently generated, enriching the video's visual materials, and the cyclic generation mechanism can meet the diverse picture quantity requirements of different scenes and needs.

[0087] Among them, the AI text-to-image engine generates high-quality pictures based on semantics, and supports intelligent cropping, scaling, and rotating of images to meet the needs of different split mirrors. In addition, a self-trained special Zilu model (based on Stable Diffusion lora training) is used to generate picture materials in batches.

[0088] As shown in Figure 3 , the content based on the Stable Diffusion lora training Zilu model includes first preparing the dataset: collecting 10-50 clear target material pictures, using the Stable Diffusion tool to uniformly crop and scale these material pictures to 512x512 resolution, automatically or manually labeling and organizing them into a text file, and creating a specific folder to store the pictures and labels.

[0089] Further, the parameters are set: configure the learning rate, network dimension, Alpha value and other key parameters for the model, enable or disable regularization as needed, add extra pictures to prevent overfitting. Then the model is trained: select the corresponding base model of the style and run the training after loading, save the model regularly. Finally, the model is tested and optimized: load the model for testing on WebUI, analyze the loss curve to evaluate performance, and iterate optimization based on the results, adjust parameters or process data and retrain.

[0090] Step seven: the text-to-speech (TTS) system adds voice and returns.

[0091] Specifically, the text-to-speech (TTS) system adds voice to the script and returns the voice file to the content generation platform. Adding voice gives the video a sound element, enhancing the video's expressiveness and appeal. The self-developed TTS system ensures natural and smooth voice and a high degree of fit with the script content.

[0092] Among them, the voice generated by the deep learning model is natural and smooth, accurately conveying emotions and information, enhancing appeal, and supporting popular voice cloning on the market, simulating the voices of well-known anchors, further increasing attention.

[0093] Step eight: the content generation platform synthesizes the video.

[0094] Specifically, the content generation platform integrates the generated pictures and voice to complete the synthesis of the video. The video synthesis process organically integrates various materials to form a complete video work, ensuring the coordination and coherence of the video in terms of picture and sound.

[0095] Among them, the video synthesis server integrates voice, pictures and external materials (such as transition animations and BGM), automatically splices audio, pictures and pre-set templates according to the timing of the shot sequence, uses multi-material fission technology, and based on the preparation of multiple materials for the same script, supports the fission production of multiple videos from the same video script, meeting the needs of different distribution channels and audience groups, and expanding the scope of dissemination.

[0096] Step nine: the user previews and downloads the video.

[0097] Specifically, the user previews the generated video and downloads the video file after confirming that it is correct. The preview function allows the user to make a final check of the video before downloading, ensuring that the video meets their requirements. Downloading the video file makes it easy for the user to share and disseminate it later.

[0098] Through the above process, the application scenario embodiment realizes the creation of personalized videos through the cooperation of multiple systems such as user operation, content generation platform, Zilu large language model, AI text-to-image, and TTS. From the user initiating the creation, through script generation and setting, scene matching, image generation, voice adding to video synthesis, and finally user preview and download, the entire process is efficient, orderly and highly flexible. Each system fully utilizes its own advantages and works collaboratively to meet the user's demand for personalized video creation, greatly improving the efficiency and quality of video creation and providing an innovative solution for the field of digital content creation.

[0099] In another application scenario, a new media company needs to quickly generate diverse short video content for different customers and distribute it on multiple social media platforms. The system includes a user terminal (Web / App), a Zilu large model server, a TTS processing engine, an AI text-to-image engine, and a video synthesis server. Each part interacts with data through the network.

[0100] Specifically, first, the staff of the new media company inputs the demand text, such as "generate a promotional video for a fashion and beauty product", through the visual interface of the user terminal (Web / App), and selects scene parameters, such as the target audience being young women and the video style being lively and fashionable.

[0101] Among them, the visual interface is simple and convenient to operate, reducing the threshold for use, so that non-professionals can easily get started. At the same time, detailed scene parameter selection helps the system to more accurately generate video content that meets the demand.

[0102] Further, the user terminal sends the input demand text and scene parameters to the Zilu large model server, which extracts keywords such as "fashion and beauty product", "young women", and "lively and fashionable" through semantic recognition, and generates a shot script by calling a pre-trained model in combination with the context, including scene description, dialogue script, etc., and outputs it in JSON format.

[0103] Among them, the Zilu large model server can quickly and accurately generate high-quality video scripts, providing a solid foundation for subsequent video production. The JSON format output facilitates data processing and transmission.

[0104] Further, the TTS processing engine receives the script text output by the Zilu large model server and converts it into WAV audio with emotion labels, such as "lively tone". Emotion-labeled voice synthesis makes the audio more appealing and can better attract the audience's attention, enhancing the spread of the video.

[0105] Further, the AI text-to-image engine generates pictures according to the shot description, such as "a beauty expert is trying a fashion makeup product", and simultaneously calls the Little Deer model to generate different pose materials of the same IP in batches, supports multi-format output, the AI text-to-image engine can quickly generate high-quality pictures, and supports batch generation, improves the production efficiency, and the application of the Little Deer model ensures the consistency and uniqueness of the picture materials, and enhances the brand recognition.

[0106] Further, the video synthesis server automatically splices the audio, pictures and preset templates (transition animation, BGM) according to the shot timing, generates five differentiated version videos by using the multi-material fission technology, such as different background colors, text positions and the like.

[0107] Among them, the automatic splicing and multi-material fission technology greatly improves the video production efficiency, and the multiple differentiated version videos generated can meet the needs of different social media platforms, and expand the video dissemination range.

[0108] Further, the user terminal (Web / App) displays the generated video, the user can preview the video and select to download or directly publish to the social media platform, so that the user can preview and select the video in time, and facilitates the distribution and promotion of the video. At the same time, the system supports directly publishing the video to the social media platform, improving the work efficiency.

[0109] Through the above process, the embodiment realizes the full-process automation from user demand input to video output and distribution through the automatic video creation and distribution based on the system architecture, fully plays the role of each component, and through the collaborative work between each part, can quickly and efficiently generate diversified short video content and meet the needs of different social media platforms, and provides strong support for the business development of new media companies.

[0110] The following is an apparatus embodiment of the present application, which can be used to execute the automatic content generation method involved in the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the automatic content generation method involved in the present application.

[0111] Please refer to Figure 4 In the embodiment of the present application, an automatic content generation device 800 is provided.

[0112] The automatic content generation device 800 includes but is not limited to a video script generation module 810, an image-to-speech conversion module 830, a style fusion unification module 850 and a video editing and synthesis module 870.

[0113] The video script generation module 810 is configured to obtain a requirement text, perform deep semantic analysis on the requirement text by using a large language model to obtain a script, and generate a video script according to the script by using a pre-trained script generation model.

[0114] The image voice conversion module 830 is configured to generate an image according to the script and the video script by performing semantic understanding by using an AI technology and an image generation algorithm, and convert the script into voice by using a text-to-speech technology and a deep learning model.

[0115] The style fusion and unification module 850 is configured to perform style analysis on the script, the image, and the voice to extract style features, and perform style fusion on the script, the image, and the voice according to the style features and style requirements in the requirement text.

[0116] The video editing and synthesis module 870 is configured to align the script, the image, and the voice in a time sequence, and edit and integrate the script, the image, and the voice by using a video synthesis technology to generate a plurality of versions of video content.

[0117] It should be noted that, in the automatic content generation provided in the above embodiments, only the division of the above functional modules is used as an example for illustration, and in actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the automatic content generation device is divided into different functional modules to complete all or part of the functions described above.

[0118] In addition, the embodiments of the automatic content generation device and the automatic content generation method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs an operation has been described in detail in the method embodiments, which will not be described herein again.

[0119] Figure 5 According to an exemplary embodiment, a structure of an electronic device is shown.

[0120] It should be noted that the electronic device is only an example adapted to the present application, and should not be considered as providing any limitation on the use range of the present application. The electronic device should also not be interpreted as needing to depend on or must have Figure 5 One or more components in the exemplary electronic device 2000 shown.

[0121] The hardware structure of the electronic device 2000 can have great differences due to different configurations or performances, such as Figure 5 As shown, the electronic device 2000 includes a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0122] Specifically, the power supply 210 is configured to provide operating voltage for each hardware device on the electronic device 2000.

[0123] The interface 230 includes at least one wired or wireless network interface 231 configured to interact with external devices. Of course, in the remaining examples of the present application, the interface 230 can further include at least one serial-parallel conversion interface 233, at least one input-output interface 235, at least one USB interface 237, and the like, as shown in the figure, which is not specifically limited herein. Figure 5

[0124] The storage 250 is a carrier for storing resources, which can be a read-only memory, a random access memory, a magnetic disk, an optical disk, or the like, and the resources stored thereon include an operating system 251, an application program 253, and data 255, and the like, and the storage mode can be temporary storage or permanent storage.

[0125] The operating system 251 is configured to manage and control each hardware device on the electronic device 2000 and the application program 253, so as to realize the operation and processing of the central processing unit 270 on the mass data 255 in the storage 250, which can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, and the like.

[0126] The application program 253 is computer readable instructions for completing at least one specific work based on the operating system 251, which can include at least one module (not shown), and each module can include computer readable instructions for the electronic device 2000. For example, the automated content generation device can be considered as an application program 253 deployed on the electronic device 2000. Figure 5

[0127] The data 255 can be signal information and the like, and is stored in the storage 250.

[0128] The central processing unit 270 can include one or more processors, and is configured to communicate with the storage 250 through at least one communication bus, so as to read the computer readable instructions stored in the storage 250, and further realize the operation and processing of the mass data 255 in the storage 250. For example, the automated content generation method is completed by the central processing unit 270 reading a series of computer readable instructions stored in the storage 250.

[0129] In addition, the present application can also be realized by hardware circuit or hardware circuit combined with software, and therefore, the realization of the present application is not limited to any specific hardware circuit, software, and combination of the two.

[0130] Please refer to Figure 6 ​​In the embodiment of the present application, an electronic device 4000 is provided, which can include a desktop computer, a notebook computer, a server, etc. with sensor identification capability.

[0131] In Figure 6 The electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0132] The data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. The communication bus 4002 can include a channel for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 In the embodiment of the present application, only one thick line is used to represent the communication bus 4002, but it does not mean that there is only one bus or only one type of bus.

[0133] Optionally, the electronic device 4000 can further include a transceiver 4004, which can be used for data interaction, such as data transmission and / or data reception, between the electronic device and other electronic devices. It should be noted that the transceiver 4004 is not limited to one in actual application, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0134] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the present disclosure. The processor 4001 can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.

[0135] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions in the form of programs or data structures and that can be accessed by the electronic device 4000, but is not limited thereto.

[0136] The memory 4003 stores computer readable instructions, and the processor 4001 can read the computer readable instructions stored in the memory 4003 through the communication bus 4002.

[0137] The computer readable instructions are executed by the one or more processors 4001 to implement the automated content generation method in the above embodiments.

[0138] In addition, the embodiment of the present application provides a storage medium, and the storage medium stores computer readable instructions, and the computer readable instructions are executed by one or more processors to implement the automated content generation method as described above.

[0139] The embodiment of the present application provides a computer program product, and the computer program product includes computer readable instructions, the computer readable instructions are stored in a storage medium, and one or more processors of an electronic device read the computer readable instructions from the storage medium, load and execute the computer readable instructions, so that the electronic device implements the automated content generation method as described above.

[0140] Compared with the related art, the present application has the following beneficial effects:

[0141] 1. The present application can significantly improve the video generation efficiency; by integrating script generation, picture generation, voice synthesis and other AI components, end-to-end automated batch video content generation is realized, the single video production time is compressed to within 10 minutes, the human participation degree is reduced by 80%, and the problem of long generation cycle caused by multi-tool cooperation and manual intervention in video creation in the prior art is solved.

[0142] 2. The application has the ability to ensure the coordination of multi-modal content; through the standardized protocol interface, the association rules of script semantics and shot list are established, ensuring that the multi-modal content (text, picture, voice) generated by AI strictly matches in time sequence and theme, eliminating the problem of manual adjustment of logical consistency among text, picture and video in traditional solutions.

[0143] 3. The application supports the scalability of the video synthesis system; through the modular architecture design, the extension capability is reserved, supporting flexible replacement and combination of video templates, AI components and adaptation rules, realizing one-time development and multi-scenario reuse, solving the problem that the existing system cannot quickly adapt to new operation scenarios and needs to redevelop functional modules.

[0144] 4. The application can generate high-quality videos containing rich elements; based on AI capabilities, picture videos containing voice, subtitles and BGM can be generated, and batch creation of short videos is realized, that is, one script can generate multiple style videos at the same time, and preview and download are supported, improving the controllability and quality of generated videos, and adapting to more types of short video fields.

[0145] 5. The application has high flexibility in video script creation; after importing company propaganda materials for creation, rewriting, expansion and contraction can be performed, and word limit and style can be selected; topic creation is also supported, meeting the video content needs in different scenarios.

[0146] 6. The application helps to achieve the traffic exposure target through rapid replication of modes; due to the support of batch generation and diversified style videos, the market demand can be quickly responded to, and diversified and high-quality video content can be generated, providing strong support for traffic exposure.

[0147] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless explicitly stated in this article, the execution of these steps has no strict sequence limitation, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0148] The above only describes some embodiments of the application, and it should be pointed out that for ordinary skilled persons in the art, without departing from the principles of the application, some improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the application.

Claims

1. An automated content generation method, characterized in that, The method includes: Obtain the requirement text, perform deep semantic analysis on the requirement text using a large language model to obtain the script, and use a pre-trained script generation model to generate a video script based on the script; Images are generated by semantic understanding of the script and video script using AI technology and image generation algorithms, and the script is converted into speech using text-to-speech technology and deep learning models. Style analysis is performed on the script, images, and audio to extract style features, and style fusion is performed on the script, images, and audio based on the style features and the style requirements in the requirement text. The script, images, and audio are aligned in chronological order, and video compositing technology is used to edit and integrate the script, images, and audio to generate multiple versions of video content.

2. The automated content generation method as described in claim 1, characterized in that, The process of using a large language model to perform deep semantic analysis on the requirement text to obtain a script, and using a pre-trained script generation model to generate a video script based on the script, includes: The large language model is used to extract key information, intent, and emotional tone from the requirement text to obtain semantic information. A pre-trained script generation model is then used to generate a structured script based on the semantic information. The requirement text includes video style, target audience, emotional tone, background environment, and character settings. The script is subjected to semantic analysis and structured processing to extract key elements and generate a video script; the key elements include scene, dialogue, rhythm, emotion, duration and shot transition information; the video script includes scene description, character dialogue and action instructions.

3. The automated content generation method as described in claim 1, characterized in that, The process of generating images based on semantic understanding of the script and video script using AI technology and image generation algorithms includes: The script and video script are semantically understood to extract key visual elements and scene descriptions. High-quality images are generated based on the visual elements and scene descriptions using AI technology and image generation algorithms. The images include versions with multiple styles and resolutions. The image generation algorithm includes generative adversarial networks.

4. The automated content generation method as described in claim 1, characterized in that, The process of converting the script into speech using text-to-speech technology and a deep learning model includes: The text content in the script is converted into natural and fluent speech using text-to-speech technology. The emotional color of the speech is optimized using a deep learning model, and the timbre, speech rate and intonation of the speech are adjusted.

5. The automated content generation method as described in claim 1, characterized in that, The step of performing style analysis on the script, images, and speech to extract style features, and then performing style fusion on the script, images, and speech based on the style features and the style requirements in the demand text, includes: The style of the script, images, and voice is analyzed to extract key style features, and style matching is performed on most images, voice, and text according to the style requirements in the required text. The images, audio, and text are fused and optimized to ensure that the script, images, and audio maintain a consistent style; the fusion optimization includes adjusting the color tone of the images, the intonation of the audio, and the layout of the text.

6. The automated content generation method as described in claim 5, characterized in that, The process of using video compositing technology to edit and integrate the script, images, and audio to generate multiple versions of video includes: The aligned images, audio, and text are edited, and video synthesis technology is used to synthesize the images, audio, and text into a video. The video is then previewed and exported in multiple formats. The editing includes adding transition effects, background music, and subtitles.

7. The automated content generation method as described in claim 1, characterized in that, The method further includes: AI technology is used to match and replace materials, add transitions and effects to generate multiple versions of the video. The generated videos are automatically detected, and those with quality problems are deleted. The quality problems include audio and video desynchronization, black frames, and frame skipping.

8. An automated content generation device, characterized in that, The device includes: The video script generation module is used to obtain the requirement text, perform deep semantic analysis on the requirement text using a large language model to obtain the script, and generate a video script based on the script using a pre-trained script generation model. The image-to-speech conversion module is used to generate images based on the script and video script through semantic understanding using AI technology and image generation algorithms, and to convert the script into speech through text-to-speech technology and deep learning models. The style fusion and unification module is used to perform style analysis on the script, images and voice to extract style features, and to perform style fusion on the script, images and voice according to the style features and the style requirements in the requirement text; The video editing and compositing module is used to align the script, images, and audio in chronological order, and to edit and integrate the script, images, and audio using video compositing technology to generate multiple versions of video content.

9. An electronic device, characterized in that, include: At least one processor and at least one memory, wherein, The memory stores computer-readable instructions; The computer-readable instructions are executed by one or more of the processors, causing the electronic device to implement the automated content generation method as described in any one of claims 1 to 7.

10. A storage medium having computer-readable instructions stored thereon, characterized in that, The computer-readable instructions are executed by one or more processors to implement the automated content generation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • SSML text automatic generation method and device fusing sub-mirror hierarchy information

    CN121072489A