Template-based intelligent video automatic synthesis method, system and device, and medium
By splitting videos into semantic segments using templates and processing audio sources in parallel to generate composite videos, the problem of low video production efficiency in vertical fields is solved, achieving efficient and professional customized video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-26
AI Technical Summary
Existing video production solutions suffer from low efficiency, difficulty in reuse, and high risk in the production of structured videos in vertical industries. They also fail to achieve integrated voice-driven video generation and are complex to operate.
A template-based intelligent video automatic synthesis method is adopted. By splitting the video into semantic segments, multiple audio sources are processed in parallel using an audio mixing pipeline, and the target engine is called for segmentation processing and resolution scaling to generate a synthesized video.
It enables batch and customized video production in vertical industries, lowers the operational threshold, ensures professional processing, solves the problem of multi-track audio processing, and avoids the security risks of cloud rendering.
Smart Images

Figure CN122293945A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video editing and multimedia content production and compositing technology, and in particular to a template-based intelligent video automatic compositing method, system, device, and medium. Background Technology
[0002] With the rapid development of video platforms, content creators have placed higher demands on the efficiency of video production, especially in vertical fields such as medical science popularization and education and training, where there is a need to mass-produce structured video content.
[0003] Currently, video production solutions on the market mainly include: 1) Traditional video editing software solutions: such as professional editing software like Adobe Premiere and Final Cut Pro. These require users to have strong technical skills, have a high learning curve, and require manual parameter adjustments for each video, resulting in low production efficiency and difficulty in meeting the needs of mass production. 2) Online video production platform solutions: such as CapCut and BitTorrent. While these platforms lower the barrier to entry, they lack template flexibility, have fixed video structures, and cannot meet the customized needs of specific industries. Furthermore, most platforms only support cloud rendering, posing risks to data security and privacy protection. 3) Procedural video generation solutions: such as using FFmpeg scripts to batch process videos. These require developers to write complex command-line scripts, lack visual preview, are difficult to debug, and lack the ability to handle complex scenarios such as audio-visual synchronization and multitrack mixing.
[0004] Existing solutions cannot break down videos into semantically meaningful functional segments (such as cover, title, main content, and ending), requiring each production to start from scratch and making it impossible to reuse existing creative designs. Although some solutions incorporate AI, TTS speech synthesis services and video editing tools belong to different systems. Users need to generate speech on an external platform first and then import it into the editing tool, which is cumbersome and prevents real-time listening and adjustment of speech effects during editing. This hinders integrated speech-driven video generation and increases operational complexity. Summary of the Invention
[0005] In view of this, the present disclosure provides a template-based intelligent video automatic synthesis method, system, device, and medium that can solve the problems of low production efficiency, difficulty in reuse, and high risk in the production of structured videos in vertical fields in the prior art.
[0006] In a first aspect, embodiments of this disclosure provide a template-based intelligent video automatic synthesis method, including: Based on the original video and user-defined cover, title, main video, and back cover templates, several segment information is automatically generated. These segment information includes cover information, title information, main video information, and back cover information, and each segment information is configured independently. The audio sources in the several segmented information are processed in parallel based on the constructed audio mixing pipeline, and the mixed audio is obtained according to the audio filter corresponding to each segmented information and the processed audio source. Determine the resolution scaling strategy for each segment; The target engine is invoked to process each segment of the segmented information and the mixed audio, and a synthesized video is automatically generated according to the resolution scaling strategy.
[0007] Secondly, embodiments of this disclosure also provide a template-based intelligent video automatic synthesis system, comprising: The splitting module is used to automatically generate several segment information based on the original video and user-defined cover segment template, title segment template, main video segment template, and back cover segment template. The several segment information includes cover information, title information, main video information, and back cover information, and each segment information is configured independently. A parallel processing module is used to process the audio sources in the several segmented information in parallel based on the constructed audio mixing pipeline; The audio mixing module is used to obtain mixed audio based on the audio filter corresponding to each segment information and the processed audio source; An automatic synthesis module is used to call the target engine to process each segment of the segmented information and the mixed audio, and automatically generate a synthesized video according to the resolution scaling strategy of each segment. The synthesized video is a video synthesized from the cover segment video, title segment video, main video segment video and back cover segment video.
[0008] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the template-based intelligent video automatic synthesis methods described above.
[0009] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions; the computer instructions are used to cause a computer to execute any of the template-based intelligent video automatic synthesis methods described above.
[0010] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0011] The template-based intelligent video automatic synthesis method provided in this disclosure first automatically generates several segment information based on the original video and user-defined cover, title, main video, and back cover templates. Multiple template information segments can be flexibly split according to user needs. Second, the audio sources in the segment information are processed in parallel based on a constructed audio mixing pipeline. Mixed audio is obtained based on the audio filter corresponding to each segment and the processed audio source. This application supports parallel processing of multiple audio sources and automatic timing alignment. Finally, the target engine is called to process each segment of the segment information and mixed audio, and the synthesized video is automatically generated based on the resolution scaling strategy of each segment. This application uses semantic segment templates as its core, splitting the video into reusable functional segments. Combined with an automated audio mixing pipeline and resolution adaptation, it balances efficiency and customization, while integrating TTS voice and video editing processes, solving the problem of multi-track audio processing, and avoiding the security risks of cloud rendering. Furthermore, it can accurately adapt to the batch and customized video production needs of vertical fields such as medical and educational institutions, reducing the operational threshold while ensuring professional processing.
[0012] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart illustrating the template-based intelligent video automatic synthesis method provided in the embodiments of this disclosure.
[0015] Figure 2This is a flowchart illustrating a method for parallel processing of audio sources in several segmented information based on a constructed audio mixing pipeline, as provided in an embodiment of this disclosure.
[0016] Figure 3 This is a flowchart illustrating a method for acquiring mixed audio provided in an embodiment of the present disclosure.
[0017] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation
[0018] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0019] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0020] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0021] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0022] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0023] Reference Figure 1 This application discloses a template-based intelligent video automatic synthesis method, including: S100 automatically generates several segment information based on the original video and user-defined cover, title, main video, and back cover templates.
[0024] The information is divided into several segments, including cover information, title information, main video information, and back cover information, and each segment is configured independently. In this embodiment, the complete video is broken down into four semantic segments: cover segment, title segment, main video segment, and back cover segment. Each segment is independently configured with materials, text, position, style, and duration, and the template structure is uniformly described using JSON Schema.
[0025] S200 processes audio sources from several segmented information in parallel based on a constructed audio mixing pipeline.
[0026] The constructed audio mixing pipeline supports parallel processing of four audio sources: background music, original video sound, TTS voice, and typing sound effects.
[0027] S300 obtains mixed audio based on the audio filter corresponding to each segment information and the processed audio source.
[0028] S400 determines the resolution scaling strategy for each segment.
[0029] The S500 calls the target engine to process each segment of information and mixed audio, and automatically generates a synthesized video based on the resolution scaling strategy.
[0030] The composite video is a video created by combining the cover video, title video, main video, and back cover video.
[0031] This application discloses a template-based intelligent automatic video synthesis method. The semantic segmentation templates are reusable, eliminating the need for creation from scratch and adapting to the batch generation needs of structured content in the medical and educational fields. Each segment template is independently configurable (cover / title / body / back cover), retaining the efficiency of template-based methods while meeting the customized needs of vertical industries. It integrates TTS speech generation, audio mixing, and video synthesis without cross-platform operation and supports real-time listening and adjustment. A unified audio mixing pipeline solves the problem of multi-track audio synchronization and mixing, and a resolution scaling strategy ensures standardized video output, requiring no professional editing skills. This application uses semantic segmentation templates as its core, breaking down videos into reusable functional segments. Combined with an automated audio mixing pipeline and resolution adaptation, it balances efficiency and customization, while also streamlining the TTS speech and video editing workflow, solving the challenges of multi-track audio processing, and mitigating the security risks of cloud rendering. Furthermore, it can accurately adapt to the batch and customized video production needs of vertical fields such as medical and educational institutions, lowering the operational threshold while ensuring professional processing.
[0032] Specifically, for S100, the following steps are taken: configuring the key data structure of each template in the front-end application layer to obtain several basic segments, including cover segment template, title segment template, main video segment template, and back cover segment template; obtaining several custom segments based on the user input received in the several basic segments; the several custom segments include user-defined cover information, title information, main video information, and back cover information.
[0033] The user-defined cover information includes the cover material type (image or video), cover material content, cover display duration, and cover background audio.
[0034] User-defined title information includes title background image, title text (including title content, font size (i.e., the font size configured on the front end), color, font weight, original title position (center, left-aligned, or right-aligned), title display duration, title word-by-word display effect (word-by-word display delay time, character interval time), and typing sound effects.
[0035] User-defined main video information includes the main video content, side banner images and their text information (including content, font size), font color, font weight, font horizontal position percentage, font vertical position percentage, original video and its audio, and other attributes.
[0036] User-defined back cover information includes the type of back cover material (image or video), the content of the back cover material, the display duration of the back cover, and the background audio of the back cover.
[0037] In this application, videos are broken down into semantic functional segments (cover / title / body / back cover), each with an independently configurable and reusable template, avoiding the need to create from scratch each time; parameters such as resolution scaling and audio mixing are automatically processed, eliminating the need for manual adjustments to each video and directly adapting to batch production needs; users do not need to master professional editing skills, only need to configure templates for automatic generation, reducing learning and operation costs.
[0038] Reference Figure 2 The S200 method, "parallel processing of audio sources in several segmented information based on a constructed audio mixing pipeline," specifically includes: S210, construct the index corresponding to each audio segment.
[0039] Specifically, the number of video segments is counted first. The background music index starts from the number of video segments, the TTS audio index is increased by 1, and the typing sound effect index is increased by 1 again. The index allocation ensures that each audio source has a unique input identifier.
[0040] S220, determine the total video duration corresponding to all segments, i.e., the sum of the display durations.
[0041] S230 determines the number of loops based on the total video duration and the duration of the preset background music.
[0042] The number of loops is the rounded-up quotient of the total video duration divided by the preset background music.
[0043] S240 uses FFmpeg to obtain the background music audio filter for the entire video based on the number of loops, total video duration, preset background music, and preset background volume.
[0044] Specifically, the aloop filter is used to achieve looping, the atrim filter is used to trim the video to the total duration, and the volume filter is used to adjust the volume, such as converting the percentage at the beginning to a decimal value between 0 and 1. If the background music duration is less than the total video duration, the system will automatically loop it; if it is longer, it will be trimmed to the total video duration.
[0045] S250: Obtain the start time of the title segment and determine the start time of the TTS voice based on the preset delay time.
[0046] Specifically, configure the timing alignment of the TTS audio, with a preset delay time preferably of 0.2 seconds. The TTS audio plays 0.2 seconds after the start of the title segment. The system first calculates the start time of the title segment (equal to the duration of the cover segment), then adds the 0.2-second delay to obtain the actual start time of the TTS. The delay is implemented using the `adelay` filter, with the delay parameter being the start time multiplied by 1000 (converted to milliseconds), and the same delay value is set for both left and right channels. The TTS volume is set using the `volume` filter (typically 0.8, or 80%).
[0047] S260 aligns the timing of typing sound effects according to the start time of the title segment.
[0048] Specifically, the typing sound effect plays immediately at the start of the title segment, synchronized with the word-by-word display. The system uses the start time of the title segment as the delay time for the sound effect, implements the delay using an adelay filter, and adjusts the sound effect volume using a volume filter based on the volume percentage configured on the front end.
[0049] S270 identifies segments containing the original audio of the video and uses the aformat filter to unify the audio format during the video standardization phase.
[0050] This step refers to the processing of the original video audio. Specifically, during the segment preprocessing stage, the system marks whether each segment contains audio. For segments containing audio, the aformat filter is used during the video normalization stage to unify the audio format (sampling rate 48000Hz, stereo); for segments not containing audio, the anullsrc filter is used to generate a silent track of corresponding duration, ensuring that all segments have an audio stream.
[0051] It should be noted that this embodiment is a preferred embodiment, in which the processing order of the four audio sources—background music, original video sound, TTS voice, and typing sound effects—is not sequential, but preferably processed in parallel, all of which are within the protection scope of this application.
[0052] This application constructs a unified audio mixing pipeline that supports parallel processing of multiple audio sources, including background music, original video audio, TTS voice, and typing sound effects. It automatically generates audio filters and amix mixing logic, eliminating the need to manually write complex scripts and solving the problem of multi-track audio mixing.
[0053] Reference Figure 3 The method of S300, which "obtains mixed audio based on the audio filter corresponding to each segment information and the processed audio source," includes the following methods for obtaining mixed audio: S310: Determine the material type corresponding to each segment, and dynamically determine the audio filter acquisition method for each segment based on the material type.
[0054] S320, obtains the audio filter corresponding to each segment in parallel according to the audio filter acquisition method.
[0055] Specifically, for the cover segment, if the cover segment's material type is video, the audio filter acquisition method is determined to be the first strategy (i.e., using FFmpeg to obtain the original audio filter for the cover segment's video). If the cover segment's material type is an image and has background music for the cover image, the audio filter acquisition method is determined to be the second strategy (i.e., using FFmpeg to obtain the background music audio filter for the cover segment).
[0056] Specifically, for the title segment, FFmpeg is used to obtain the audio filter corresponding to the typing sound effect of the title segment; TTS speech synthesis technology is used to convert the title content into target speech, and FFmpeg is used to obtain the audio filter corresponding to the target speech; Specifically, for the main video segment, FFmpeg is used to obtain the original video and audio filters of the main video segment; FFmpeg is also used to process the side banner images and their text information to generate image filters corresponding to the side banner images.
[0057] Specifically, for the back cover section, if the material type of the back cover section is video, the audio filter acquisition method is determined to be the first strategy (i.e., using FFmpeg to obtain the original audio filter of the video in the back cover section); if the material type of the back cover section is an image and has background music for the back cover image, the audio filter acquisition method is determined to be the second strategy (i.e., using FFmpeg to obtain the background music audio filter of the back cover section).
[0058] S330 processes the corresponding audio source based on the audio filter corresponding to each segment to obtain the optimized audio source corresponding to each segment.
[0059] Specifically, this includes background music looping, TTS delay, and sound effect delay.
[0060] S340 determines the corresponding mixing strategy based on the number of optimized audio sources and generates mixed audio based on the mixing strategy.
[0061] Specifically, collect tags for all optimized audio sources. If the number of audio sources is greater than 1, generate an amix blending filter; if it is equal to 1, directly copy the audio source; if it is equal to 0, generate a silent track.
[0062] In existing technologies, TTS speech generation and video editing belong to different systems, requiring manual import of speech and making real-time adjustments impossible. This application integrates TTS speech synthesis into the audio mixing pipeline, using it as one of the audio sources for parallel processing with other audio, achieving integrated speech generation, audio mixing, and video synthesis. During the template configuration stage, TTS speech filters (such as delay and volume) can be adjusted in real time, and the matching effect of speech and video segments can be previewed simultaneously, eliminating the need for cross-platform operations and simplifying the process.
[0063] Furthermore, FFmpeg can be used to process the original audio filters of the cover segment, or the background music audio filters of the cover segment, the audio filters corresponding to the typing sound effects of the title segment, the original audio filters of the main video segment, the original audio filters of the back cover segment, or the background music audio filters of the back cover segment, the background music audio filters of the entire video, and the audio filters corresponding to the target speech to obtain mixed audio.
[0064] The method for S400 to "determine the resolution scaling strategy for each segment" specifically includes: determining the front-end preview resolution and the back-end output resolution; determining the global scaling factor based on the front-end preview resolution and the back-end output resolution; determining the back-end rendered font size based on the global scaling factor; and determining the back-end output position information based on the front-end position parameters and the global scaling factor.
[0065] The global scaling factor is calculated as the backend output resolution divided by the frontend preview resolution, with the preferred frontend preview resolution being 960×540.
[0066] For the cover section, the output resolution of the back end is preferably 1920 (width) × 1080 (height); for the back cover section, it is preferably the same as the cover section.
[0067] For the title segment, obtain the actual resolution of the title background image; based on the actual resolution and the backend output resolution, determine the width scaling ratio and the height scaling ratio, and take the smaller value of the width scaling ratio and the height scaling ratio as the scaling factor of the title segment; based on the scaling factor and the font size configured on the frontend, obtain the backend rendered font size corresponding to the title segment, that is, the backend rendered font size corresponding to the title segment is the font size configured on the frontend multiplied by the scaling factor.
[0068] For the main video segment: determine the resolution of the main video segment (preferably 1920×1080 for backend output); obtain the backend rendering font size corresponding to the main video segment based on the target scaling factor and the font size configured on the frontend (i.e., the font size configured on the frontend multiplied by the scaling factor); determine the horizontal and vertical pixel positions of the backend output corresponding to the main video segment based on the horizontal position percentage, the vertical position percentage, and the resolution of the main video segment.
[0069] Specifically, the front-end uses a percentage coordinate system (0-100). During back-end rendering, the percentage is first converted into pixel values at the preview resolution, and then multiplied by a scaling factor to obtain the pixel values at the output resolution. For example, if the front-end is configured with a horizontal position of 50% (i.e., 50% of the preview width), the back-end calculates it as 960 × 0.5 × 2 = 960 pixels (50% of the output width).
[0070] Furthermore, for the color format of each segment, if the front-end uses RGBA format (such as rgba(14,14,14,1)), the back-end needs to convert it to a hexadecimal format supported by FFmpeg (such as #0E0E0E). The system extracts the RGB components using regular expressions and converts them to a hexadecimal format supported by FFmpeg.
[0071] Furthermore, if the user defines a requirement for the title to be displayed character by character, this application also includes: 1) splitting the title text into several character groups (each group contains one character) to obtain the cumulative text corresponding to each character position; 2) determining the initial width of the text (i.e., the width of a single character × the total number of single characters) based on the number of single characters in the title text; 3) determining the starting position of the first character based on the initial width of the text and the title position.
[0072] To prevent text position from shifting when displaying each character, the system pre-calculates the initial width of the complete text. When the title is centered, the starting position of the first character = width of the back-end output resolution (1920) ÷ 2 - initial text width ÷ 2; when the title is left-aligned or right-aligned, the starting position of the first character = original title position × scaling factor.
[0073] To split the title text into several character groups, for example, for the nth character, construct a cumulative text string containing the first n characters. For example, the title "First Line" is split into ["First", "One", "Line"], where the cumulative text of the first character is "First", the cumulative text of the second character is "First", and the cumulative text of the third character is "First Line".
[0074] This application also includes: constructing a timing control expression, specifically determining the start and end times of the cumulative text corresponding to each character position (start time = (character-by-character display delay time (in milliseconds) + character position × character interval time) ÷ 1000 (in seconds); the end time is equal to the start time of the next character, and no end time is set for the last character. The start and end times are converted into an enable expression, in the format "gte(t, start time)". "lt(t, end time)" means that the text layer is displayed only within the specified time range.
[0075] According to the start time, end time, starting position of the first character, and cumulative text corresponding to all character positions, use FFmpeg to generate a text filter chain for the title segment, where each character position corresponds to a filter. Specifically, connect all the drawtext filters of all characters in sequence, separated by commas, to form a complete filter chain. Each filter contains parameters such as text content, font file path, font size, font color, X coordinate, Y coordinate, enable expression, etc.
[0076] In a preferred embodiment, a 10-character title will generate 10 drawtext filters. The first filter displays "第" from 0 to 100 milliseconds, the second filter displays "第一" from 100 to 200 milliseconds, and so on, to achieve the typewriter word-by-word display effect.
[0077] For the method of S500 "call the target engine to process each segment of several segment information and mixed audio, and automatically generate a synthesized video according to the resolution scaling strategy", it specifically includes: for each video segment, use the scale filter to scale the video to the output resolution (1920×1080), and set the scaling mode to force_original_aspect_ratio=decrease, which means to maintain the original aspect ratio. If the original video aspect ratio is inconsistent with the output resolution, the video is scaled down to fit the output resolution, and the insufficient part is filled with black. Then use the pad filter to fill the video to the exact output resolution, with the filling position centered and the filling color being black. Use the fps filter to unify the frame rate of all video segments to 25fps, ensure that the frame rate is consistent when the segments are connected, and avoid playback stuttering. Use the concat filter to connect all the standardized video segments and audio segments. For video connection, use the v=1:a=0 mode of the concat filter, which means only connect the video stream; for audio connection, use the v=0:a=1 mode of the concat filter, which means only connect the audio stream. The number of connections is equal to the number of segments. Assign labels to each input in the format of "segment index: stream type" (such as [0:v] represents the video stream of the 0th input), then generate standardized filters for each segment, with output labels in the format of [v0][a0], etc. Finally, connect all the output labels to generate the final [v] and [a] labels.
[0078] Furthermore, for TTS speech synthesis integration, the specific steps include: the front-end calls the Baidu Intelligent Speech API (or other intelligent speech APIs) to convert the title text into speech (specifically, in which step). The system first obtains a Baidu AI access token, which is valid for 30 days. The system caches the token and automatically refreshes it before expiration. Then, it constructs API request parameters, including the access token, text content, voice ID (0-111, corresponding to different characters such as Du Xiaomei, Du Xiaoyu, etc.), speech rate level (1-9, 1 being the slowest and 9 the fastest), pitch (fixed at 5), volume (fixed at 9, i.e., high volume), audio format (WAV format, encoded in 3), client identifier, and language (Chinese). After sending the POST request, it receives the returned audio data (ArrayBuffer format).
[0079] The frontend converts the generated audio data into an AudioBuffer object for real-time playback preview. Simultaneously, it converts the AudioBuffer into a Base64 encoded string and stores it in the audioBase64 field of the template's TTS object for transmission to the backend.
[0080] After receiving the template data, the backend checks the TTS configuration in the title section. If the enable flag is true and the audioBase64 field exists, TTS audio processing is performed. The system decodes the Base64 string into binary data, generates a unique temporary filename (in the format tts_audio_timestamp.wav), writes the binary data to the temporary file, and returns the file path.
[0081] During the video compositing stage, the system adds the TTS audio file as independent input to the FFmpeg command, using the adelay filter to achieve timing alignment with the title segment. The start time of the TTS audio is equal to the cover segment duration plus a 0.2-second delay, ensuring it plays after the title segment begins, coordinating with the word-by-word presentation. After video compositing is complete, the system automatically deletes the temporary TTS files, freeing up storage space.
[0082] For video segment overlay image and text processing, the specific steps include: Overlay image processing: The main video segment supports adding overlay images (such as left / right overlays). First, the overlay image is added as a separate input to the FFmpeg command. Then, the scale filter is used to scale the overlay image to the video size, with the scaling mode consistent with the video scaling. The overlay filter is then used to overlay the overlay image onto the video, with the default overlay position being the top left corner (0:0). PNG images with transparent backgrounds are supported.
[0083] If images exist, the video is scaled to the output resolution and then overlaid to cover the images. Finally, text elements are overlaid one by one. The output of each processing step serves as the input for the next step, forming a chain-like filter structure. The final output is labeled [final], used for connecting subsequent paragraphs.
[0084] Furthermore, the application also includes: supporting users to edit video templates through a visual interface by dragging, configuring, and other methods.
[0085] This application also includes: the ability to render the effects of each paragraph in real time via a paragraph previewer, supporting both paragraph-by-paragraph preview and overall preview.
[0086] This application also includes: a TTS configurator that can be configured to integrate speech synthesis services, supporting voice selection, speech rate adjustment, real-time listening, etc.
[0087] This application also includes: support for batch downloading video, image, audio and other material files from remote URLs, and support for concurrent downloads and resume interrupted downloads.
[0088] The application also includes support for batch file downloads, with each download task containing information such as a unique identifier key, file URL, and file type.
[0089] The application also includes support for a local caching mechanism. Specifically, the system first checks if a file with the same name already exists in the temporary directory. If it does, the local file is used directly to avoid duplicate downloads. If the file does not exist, it is downloaded from a remote URL, saved to the temporary directory, and its original filename is used.
[0090] In this application, when parsing the template, the system can use placeholders (in the format PLACEHOLDER_key_name) to mark the files that need to be downloaded. After the download is complete, the system iterates through the template configuration and replaces all placeholders with the actual file paths. This design decouples the template configuration from the file download, improving the system's flexibility.
[0091] Furthermore, this application also includes generating a unique temporary directory for each rendering task, in which all intermediate files (such as processed video, TTS audio, etc.) are stored; after rendering is completed, the system can choose to retain or clean up the temporary files to facilitate debugging and troubleshooting.
[0092] Secondly, this application discloses a template-based intelligent video automatic synthesis system for implementing the template-based intelligent video automatic synthesis scheme disclosed in the first aspect of this application. Specifically, the system includes: The splitting module is used to automatically generate several segment information based on the original video and user-defined cover segment template, title segment template, main video segment template, and back cover segment template. The segment information includes cover information, title information, main video information, and back cover information, and each segment information is configured independently. The parallel processing module is used to process audio sources in several segments of information in parallel based on the constructed audio mixing pipeline; The audio mixing module is used to obtain mixed audio based on the audio filter corresponding to each segment information and the processed audio source; The automatic compositing module calls the target engine to process several segmented information and mixed audio for each segment, and automatically generates a composite video based on the resolution scaling strategy determined for each segment. The composite video is a video composed of the cover segment video, title segment video, main video segment video, and back cover segment video.
[0093] Specifically, the front-end application layer in this application includes four core modules: a template editor, a paragraph editor previewer, a TTS configurator, and a rendering controller. The template editor provides a visual interface, allowing users to edit video templates through drag-and-drop and configuration. The paragraph previewer renders the effects of each paragraph in real time, supporting both segment-by-segment and overall previews. The TTS configurator integrates a speech synthesis service, supporting voice selection, speech rate adjustment, and real-time listening. The rendering controller is responsible for serializing template data into JSON format and transmitting it to the backend via the HTTP protocol.
[0094] The backend service layer in this application comprises four core modules: file download service, TTS audio service, video rendering service, and output manager. The file download service is responsible for batch downloading video, image, and audio files from remote URLs, supporting concurrent downloads and breakpoint resumption. The TTS audio service receives Base64 encoded audio data transmitted from the front end, decodes it, and saves it as a temporary WAV file. The video rendering service is a core module responsible for calling the FFmpeg engine for video compositing. The output manager is responsible for file path management, temporary file cleanup, and access URL generation.
[0095] This application also specifically includes a rendering engine layer for an FFmpeg-based multimedia processing engine, which implements complex audio and video processing through the filter_complex mechanism, including functions such as video scaling, filling, frame rate unification, audio formatting, and multitrack mixing.
[0096] The template-based intelligent video automatic synthesis solution disclosed in this application allows for visual preview and editing on the front end, and high-quality video rendering on the back end. It ensures a WYSIWYG (What You See Is What You Get) experience through a unified resolution scaling factor. A unified audio mixing pipeline supports parallel processing and automatic timing alignment of background music, TTS (Text-to-Speech), typing sound effects, and original video audio. By defining a resolution scaling strategy for each segment, all size parameters can be automatically scaled during backend rendering, effectively solving the mismatch problems in font size, element position, etc., caused by directly using front-end configuration parameters in existing technologies.
[0097] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0098] The processor may be a central processing unit (CPU) or other processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the template-based intelligent video automatic synthesis method described in the foregoing embodiments of this disclosure.
[0099] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.
[0100] like Figure 4 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 4 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0101] like Figure 4As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0102] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 4 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
[0103] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the template-based intelligent video automatic synthesis method of embodiments of this disclosure are performed.
[0104] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0105] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the template-based intelligent video automatic synthesis method described in the foregoing embodiments of the present disclosure are performed.
[0106] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).
[0107] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0108] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0109] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.
[0110] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.
[0111] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0112] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0113] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0114] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A template-based intelligent video automatic synthesis method, characterized in that, include: Based on the original video and user-defined cover, title, main video, and back cover templates, several segment information is automatically generated. These segment information includes cover information, title information, main video information, and back cover information, and each segment information is configured independently. The audio sources in the several segmented information are processed in parallel based on the constructed audio mixing pipeline, and the mixed audio is obtained according to the audio filter corresponding to each segmented information and the processed audio source. Determine the resolution scaling strategy for each segment; The target engine is invoked to process each segment of the segmented information and the mixed audio, and a synthesized video is automatically generated according to the resolution scaling strategy.
2. The template-based intelligent video automatic synthesis method according to claim 1, characterized in that, The system automatically generates several segment information based on the original video and user-defined cover, title, main video, and back cover templates, including: Configure the key data structure of each template in the front-end application layer to obtain several basic segments; the several basic segments include cover segment template, title segment template, main video segment template, and back cover segment template; Based on the user input received from the aforementioned basic segments, several custom segments are obtained; the aforementioned custom segments include user-defined cover information, title information, main video information, and back cover information.
3. The template-based intelligent video automatic synthesis method according to claim 1, characterized in that, The constructed audio mixing pipeline processes the audio sources in the segmented information in parallel, including: Build an index for each audio segment; Determine the total video duration corresponding to all segments; The number of loops is determined based on the total video duration and the duration of the preset background music. Based on the number of loops, the total video duration, the preset background music, and the preset background volume, FFmpeg is used to obtain the background music audio filter for the entire video. Obtain the start time of the title segment and determine the start time of the TTS voice based on the preset delay time; Align the timing of the typing sound effects according to the start time of the title segment; Identify the segments containing the original audio from the video and use the aformat filter to unify the audio format during the video standardization phase.
4. The template-based intelligent video automatic synthesis method according to claim 3, characterized in that, The step of obtaining the mixed audio based on the audio filter corresponding to each segment information and the processed audio source includes: Determine the material type corresponding to each segment, and dynamically determine the audio filter acquisition method for each segment based on the material type; The audio filter corresponding to each segment is obtained in parallel according to the audio filter acquisition method described above; The corresponding audio source is processed based on the audio filter corresponding to each segment to obtain the optimized audio source corresponding to each segment. A corresponding mixing strategy is determined based on the number of optimized audio sources, and mixed audio is generated based on the mixing strategy.
5. The template-based intelligent video automatic synthesis method according to claim 4, characterized in that, The process of determining the material type corresponding to each segment and dynamically determining the audio filter acquisition method for each segment based on the material type includes: If the cover image is a video, then determining the method of obtaining the audio filter is the primary strategy. If the cover image is an image and has background music, then the audio filter acquisition method is determined to be the second strategy. Use FFmpeg to obtain the audio filter corresponding to the typing sound effect of the title segment; The title content is converted into target speech using TTS speech synthesis technology, and the audio filter corresponding to the target speech is obtained using FFmpeg. The original video and audio filters for the main video segment are obtained using FFmpeg; the side banner images and their text information are processed using FFmpeg to generate image filters corresponding to the side banner images. If the material type of the back cover segment is video, the method of obtaining the audio filter is determined to be the first strategy; if the material type of the back cover segment is image and has background music for the back cover image, the method of obtaining the audio filter is determined to be the second strategy.
6. The template-based intelligent video automatic synthesis method according to claim 1, characterized in that, The determination of the resolution scaling strategy for each segment includes: Determine the front-end preview resolution and the back-end output resolution; The global scaling factor is determined based on the front-end preview resolution and the back-end output resolution. The font size for backend rendering is determined based on the global scaling factor. The backend output position information is determined based on the front-end position parameters and the global scaling factor.
7. A template-based intelligent video automatic synthesis system, characterized in that, include: The splitting module is used to automatically generate several segment information based on the original video and user-defined cover segment template, title segment template, main video segment template, and back cover segment template. The several segment information includes cover information, title information, main video information, and back cover information, and each segment information is configured independently. A parallel processing module is used to process the audio sources in the several segmented information in parallel based on the constructed audio mixing pipeline; The audio mixing module is used to obtain mixed audio based on the audio filter corresponding to each segment information and the processed audio source; An automatic synthesis module is used to call the target engine to process each segment of the segmented information and the mixed audio, and automatically generate a synthesized video according to the resolution scaling strategy of each segment. The synthesized video is a video synthesized from the cover segment video, title segment video, main video segment video and back cover segment video.
8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the template-based intelligent video automatic synthesis method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions; the computer instructions are used to cause the computer to perform the template-based intelligent video automatic synthesis method as described in any one of claims 1-6.
10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1-6.