Video mixing system
By standardizing video footage, performing semantic analysis, and automating editing, the problems of poor format compatibility, low efficiency, and insufficient terminal adaptability in video mashup technology have been solved, achieving high-quality and diverse video editing effects and meeting the needs of large-scale personalized content production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video montage technology lacks the ability to understand the semantics of multi-source video materials, resulting in insufficient narrative coherence and visual expressiveness in the editing results. It also suffers from poor format compatibility, low editing efficiency, and insufficient terminal adaptability, making it difficult to meet the needs of large-scale and personalized content production.
The material preprocessing unit standardizes the format, resolution, and quality; the intelligent editing decision unit generates editing logic based on semantic analysis; the automated editing execution unit performs efficient editing; and the effect optimization output unit enhances the visuals and optimizes the format to generate the target mashup video.
It achieves unified and standardized processing of multi-source video materials, improves video quality and stability, enhances the narrative coherence and visual expressiveness of the editing results, improves editing efficiency, and ensures that the video presents the best effect on different terminals, meeting the needs of large-scale personalized content production.
Smart Images

Figure CN120897100B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more particularly to a video editing system. Background Technology
[0002] With the explosive growth in demand for short video content creation, e-commerce platforms, self-media operations, and brand promotion scenarios have placed higher demands on the automation and intelligence of video mashup editing. As a key technology for processing multi-source video footage into a complete work, the efficiency and quality of video mashup editing directly impact the cost of content production and its dissemination effectiveness. Currently, video mashup editing solutions on the market mainly rely on manual editing or semi-automated tools based on fixed templates, which are insufficient to meet the needs of rapid processing of massive amounts of footage and personalized creation.
[0003] Existing video editing technologies typically employ a "template matching + simple parameter adjustment" approach: splicing together footage using preset editing templates, combined with manually specified transition effects and duration allocations. This approach often lacks semantic understanding of the video content and cannot intelligently process the footage based on its visual characteristics, emotional attributes, and narrative logic. For example, when faced with video footage from different scenarios (such as e-commerce product displays and brand story videos), current technology struggles to automatically identify key content and match appropriate editing strategies, resulting in significant deficiencies in narrative coherence and visual appeal.
[0004] Further analysis reveals the following technical shortcomings in existing solutions: First, the preprocessing of materials lacks a standardized mechanism, resulting in poor compatibility with materials of various formats, resolutions, and qualities, easily leading to image quality loss or format incompatibility issues in the edited video. Second, the editing decision-making process relies on human experience or fixed rules, failing to generate dynamic editing logic based on the semantic information of the materials (such as scene type and object recognition results), making it difficult to adapt to diverse creative needs. Third, during automated editing execution, steps such as shot segmentation, transition generation, and audio fusion lack intelligent optimization, resulting in low editing efficiency and monotonous effects. Fourth, the output stage lacks multi-terminal adaptation and quality optimization mechanisms, failing to perform targeted optimization based on the characteristics of the playback platform, affecting the final playback effect. These technical bottlenecks make it difficult for existing video mashup solutions to balance efficiency and quality when facing large-scale, personalized content production demands. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a video mixing and editing system to at least partially solve the above-mentioned problems.
[0006] A video editing system comprising:
[0007] The material preprocessing unit is used to preprocess the original video material to generate standardized editing material;
[0008] The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions;
[0009] The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video;
[0010] The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0012] Figure 1 This is a schematic diagram of the structure of a video mixing and editing system according to an embodiment of this application.
[0013] Figure 2 This invention provides a schematic diagram of the structure of an electronic device. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.
[0015] It should be understood that the terms "first," "second," and "third," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or sets thereof.
[0016] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any and all combinations of one or more of the associated listed items, and includes such combinations.
[0017] Figure 1 This is a schematic diagram of the structure of a video mixing and editing system according to an embodiment of this application. Figure 1 As shown, a video editing system includes:
[0018] The material preprocessing unit is used to preprocess the original video material to generate standardized editing material;
[0019] The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions;
[0020] The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video;
[0021] The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video.
[0022] The video mixing and editing system of this application has the following technical advantages:
[0023] 1. By standardizing the original video footage through the material preprocessing unit, the compatibility issues of multiple source formats, resolutions, and qualities can be effectively resolved, avoiding image quality loss caused by format incompatibility, providing a unified and standardized material foundation for subsequent editing, and improving the overall quality and stability of the video.
[0024] 2. The intelligent editing decision unit generates editing logic based on semantic analysis of standardized materials, breaking through the limitations of traditional reliance on human experience or fixed rules. The system can automatically identify semantic information such as material scene type and object characteristics, dynamically generate appropriate editing strategies, significantly enhance the narrative coherence and visual expressiveness of the editing results, and meet diverse creative needs.
[0025] 3. The automated editing execution unit performs automated processing based on intelligently generated editing decision instructions, realizing intelligent operation in shot segmentation, transition generation, and audio fusion. Compared with existing editing methods that lack intelligent optimization, it greatly improves editing efficiency, while avoiding the problem of monotonous effects and achieving high-quality and diversified editing effects.
[0026] 4. The output unit optimizes the visuals and formats of the initial montage video. Combined with multi-terminal adaptation technology, it can make targeted adjustments according to the characteristics of different playback platforms to ensure that the video presents the best playback effect on various terminals. This solves the problem of insufficient adaptability in the output stage of the existing solution, effectively balances efficiency and quality, and meets the needs of large-scale and personalized content production.
[0027] Optionally, the material preprocessing unit is used to preprocess the original video material to generate standardized edit material, specifically including the following steps:
[0028] The multi-source format standardization subunit is used to perform format-unified conversion on the original video footage to generate format-standardized footage;
[0029] The spatiotemporal resolution unification subunit is used to normalize the temporal frame rate and spatial resolution of format-standardized materials in order to generate spatiotemporally standardized materials.
[0030] The content noise filtering subunit is used to perform noise removal and abnormal frame filtering on spatiotemporally normalized material to generate denoised material.
[0031] The metadata semantic annotation subunit is used to parse and semantically annotate metadata such as shooting parameters and scene information of the denoised material to generate semantically annotated material.
[0032] Optionally, the multi-source format standardization subunit includes:
[0033] The format detection module is used to detect the format type and encoding protocol of the original video footage in order to generate format description information;
[0034] The format conversion execution module is used to perform unified conversion processing of container format and encoding protocol according to the format description information in order to generate intermediate format conversion materials;
[0035] The encoding optimization module is used to control the bitrate and optimize the encoding parameters of intermediate materials during format conversion in order to generate encoded optimized materials.
[0036] The format verification module is used to verify the format integrity and protocol compatibility of the encoded optimized materials in order to generate standardized materials.
[0037] Specifically, the format detection module employs a dual-track detection strategy of "metadata priority + deep learning assistance." First, it reads the header metadata of the video file, such as the ftyp (file type) and moov (movie structure) atoms in MP4 format, and the RIFF (Resource Interchange File Format) block structure in AVI format, extracting the container type (e.g., MP4, AVI, MOV) and basic encoding information (e.g., H.264, H.265). If the metadata is missing or corrupted, a deep learning model is activated: the first 5 seconds of the video are converted into a fixed-size image sequence and input into a pre-trained convolutional neural network (e.g., a ResNet variant). This network learns from a large amount of labeled data (containing over 100 format samples) to identify encoding feature patterns in video frames (e.g., H.264 macroblock structure, VP9 intra-frame prediction patterns), thereby determining the encoding protocol type. Finally, the container and encoding information are integrated into structured format description information, such as `{"container":"MP4","codec":"H.265"}`.
[0038] The format conversion execution module is based on an adaptive transcoding framework, and its core consists of three parts:
[0039] Intelligent parameter configuration: After receiving the format description information, the system searches the historical transcoding database and uses a similarity matching algorithm to find the optimal parameter combination for converting similar formats. For example, if the input is MOV (ProRes encoding) to MP4 (H.264 encoding), the system will prioritize selecting preset H.264 encoding parameters (such as CRF value and keyframe interval) that have previously yielded high-quality images and small file sizes.
[0040] Parallel processing engine: The video is divided into multiple segments along the timeline (e.g., every 10 seconds), and the parallel computing power of the GPU is used to process the transcoding tasks of different segments simultaneously. At the same time, pipeline technology is used so that while segment A is undergoing format conversion, segment B is simultaneously preprocessed (e.g., resolution scaling), improving overall processing efficiency;
[0041] Real-time quality monitoring: During transcoding, after each frame is processed, a lightweight image quality assessment model (such as a PSNR predictor based on a shallow neural network) is used to compare the differences between the original frame and the converted frame. If the image quality loss exceeds a threshold (e.g., 10%), the process automatically rolls back, adjusts the encoding parameters, and reprocesses the segment. The final output contains intermediate footage with basic format conversion.
[0042] The encoding optimization module prioritizes visual focus and implements intelligent encoding through three steps:
[0043] Salience region identification: The algorithm utilizes an improved U-Net network to analyze intermediate footage frame by frame, combining this with an attention mechanism to highlight key visual regions such as faces and product subjects. For example, in beauty videos, the algorithm automatically identifies the areas in the frame where products like lipstick and foundation are located.
[0044] Dynamic bitrate allocation: Higher bitrate is allocated to prominent areas to ensure clear details (such as individual strands of hair or product textures); lower bitrate is allocated to non-critical areas such as the background to reduce file size. For example, if prominent areas occupy 30% of the image, then 60% of the total bitrate is allocated to them.
[0045] Parameter search and optimization: The simulated annealing algorithm is used to iteratively search for the optimal combination within a preset encoding parameter space (e.g., CRF value range 18-32, number of B-frames 0-3). After each parameter adjustment, a 10-second test clip is generated through fast encoding, and the visual similarity is compared with the original footage until the best parameter configuration is found, at which point the encoded and optimized footage is output.
[0046] The format verification module uses triple verification of "structural protocol compatibility" to ensure material standardization:
[0047] Container structure verification: Use a state machine model to parse the binary structure of the file, check whether the atomic order of the container conforms to the standard (e.g., the moov atom of MP4 should be at the beginning or end of the file), and verify the integrity of key structural fields (e.g., timestamp, index table).
[0048] Encoding protocol compliance: Compare encoding parameters with industry standards, such as verifying whether the H.264 profile (e.g., Main Profile, High Profile) and level (e.g., Level 4.1) match the video resolution and frame rate, and detecting the presence of illegal parameters (e.g., excessively high bitrate limits);
[0049] Terminal compatibility testing: The decoding environment of mainstream playback terminals (such as TikTok app, Chrome browser, and smart TVs) is simulated, and the footage is run using sandbox technology. If playback stutters, audio-visual asynchrony, or decoding failures occur, the footage is marked as incompatible. Only when all three verifications pass is the output format standardized footage generated; otherwise, error messages are fed back to the format conversion execution module for correction.
[0050] The format description information output by the format detection module serves as a global configuration parameter, driving the format conversion execution module to select a transcoding strategy. The converted intermediate material undergoes further bitrate and parameter adjustments by the encoding optimization module to improve quality. Finally, the format verification module performs a comprehensive check on the optimization results, forming a complete chain of "configuration detection → conversion execution → optimization enhancement → verification output." If verification fails, the error message will trigger parameter rollback and reprocessing in the conversion module, ensuring that the final output material conforms to a unified standard and can seamlessly integrate with subsequent editing workflows.
[0051] Optionally, the spatiotemporal resolution unification subunit includes:
[0052] The resolution analysis module is used to perform width, resolution, and pixel ratio detection on standardized format materials to generate resolution metadata.
[0053] Extract the width, height, and pixel aspect ratio (PAR) of the video stream, and identify abnormal resolutions (such as non-standard 720×480).
[0054] The frame rate synchronization module is used to perform time frame rate detection and unified synchronization processing on resolution metadata in order to generate frame rate standardized materials.
[0055] Frame interpolation (such as motion compensated frame interpolation MCFI) is used to unify non-standard frame rates (such as 29.97fps) to 30fps, maintaining smooth motion.
[0056] The size scaling module is used to perform spatial resolution scaling and aspect ratio preservation processing on frame rate normalized footage to generate size normalized footage;
[0057] The spatiotemporal consistency verification module is used to verify the frame temporal continuity and resolution consistency of size-standardized materials in order to generate spatiotemporal standardized materials.
[0058] For the resolution parsing module, a dual-process parsing strategy of "metadata reading + pattern matching" is adopted. First, the width, height, and pixel aspect ratio (PAR) information are directly extracted from the metadata of the video stream of standardized format materials. For example, for MP4 format, the resolution data is obtained by parsing the width and height fields in the stbl (sample table) atoms, and the PAR value is derived from the display aspect ratio field. If the metadata is missing or there are conflicts, the pattern matching mechanism is activated: edge detection (such as Canny operator) is performed on the video frames to identify the geometric structures in the picture (such as rectangular borders, human outlines), and similarity calculation is performed by combining the preset standard resolution templates (such as 1920×1080, 1280×720) to determine the actual resolution. At the same time, resolution data that does not conform to the standard is marked by a preset abnormal resolution rule library (such as 720×480, non-uniform scaling size). Finally, the resolution metadata containing the original resolution, PAR value, and abnormal status is integrated, such as `{"width":1920,"height":1080,"PAR":1.0,"is_abnormal":false}`.
[0059] For the frame rate synchronization module, this module uses Motion Compensated Frame Interpolation (MCFI) technology as its core to achieve precise frame rate uniformity. First, it extracts raw frame rate information from resolution metadata. If a non-standard frame rate (e.g., 29.97fps, 59.94fps) is detected, the MCFI algorithm is activated: it analyzes the pixel motion trajectories of adjacent frames using optical flow to calculate the object's displacement vector (e.g., the direction and speed of a character's movement), and generates an intermediate transition frame based on this. For example, when converting 29.97fps to 30fps, one compensation frame is inserted every 100 frames. The pixel values of the compensation frame are generated through weighted fusion of preceding and following frames (e.g., 40% from the preceding frame and 60% from the following frame) and motion vector correction, ensuring smooth, stutter-free motion. Simultaneously, temporal filtering technology is used to eliminate ghosting or blurring that may occur during interpolation, ultimately outputting standardized footage with a uniform frame rate of 30fps or 60fps.
[0060] For the size scaling module, this module achieves resolution scaling while maintaining aspect ratio based on a "perception-first + intelligent cropping" strategy. First, the scaled size is calculated based on the target resolution (e.g., 1080P, 720P) and the original video's aspect ratio. If the original video has a non-standard aspect ratio (e.g., 2.35:1 cinematic aspect ratio), two processing methods are used: For videos that need to adapt to vertical screen platforms (e.g., TikTok), center cropping is prioritized to retain the main subject (e.g., people, products) before scaling to the target size; for scenes that need to maintain a full aspect ratio, black borders are added top, bottom, left, or right (letterboxing / pillarboxing) to ensure the aspect ratio remains unchanged. During scaling, the Lanczos interpolation algorithm is used to resample pixels, and multi-scale analysis preserves high-frequency details (e.g., text edges, textures) to avoid jagged edges or blurring caused by scaling. The final output is standardized material with a standard size and complete main subject.
[0061] For the spatiotemporal consistency verification module, it ensures the spatiotemporal continuity of the footage through a "temporal analysis + multi-frame comparison" mechanism. First, it uses inter-frame difference to detect the temporal continuity of the video: calculating the pixel difference between adjacent frames. If the difference between multiple consecutive frames exceeds a threshold (e.g., a 50% pixel change rate), it is marked as a temporal jump (e.g., dropped frames, incorrect splicing). Simultaneously, it checks the resolution consistency of the video frame-by-frame, comparing the width, height, and pixel layout of each frame using a hash algorithm to ensure there are no sudden resolution changes. For videos containing dynamic elements, it uses optical flow field analysis to smooth the trajectory of moving objects. If trajectory breaks or abnormal acceleration are detected, it is determined to be spatiotemporally inconsistent. Furthermore, the module verifies the synchronization of audio and video by comparing the timestamps of audio and video frames to ensure that the audio-visual deviation is within ±50ms. Only when all verification items pass is the spatiotemporally standardized footage output; if problems exist, they are fed back to upstream modules (e.g., frame rate synchronization, size scaling) for correction.
[0062] The metadata output by the resolution parsing module serves as the basic configuration, driving the frame rate synchronization module to select the interpolation strategy and target frame rate. The frame rate-standardized footage then enters the size scaling module, where it undergoes adaptive processing based on aspect ratio and target resolution. Finally, the spatiotemporal consistency verification module performs comprehensive validation of the output results, forming a complete chain of "configuration parsing → frame rate synchronization → size adaptation → consistency verification." If the verification detects timing or resolution anomalies, the error message will trigger parameter adjustments and reprocessing in the corresponding module, ensuring that the final footage meets standardization requirements in both temporal and spatial dimensions, providing high-quality input for subsequent editing processes.
[0063] Optionally, the content noise filtering subunit includes:
[0064] The noise detection module is used to perform multi-frame joint detection of luminance noise and color noise on spatiotemporally standardized materials to generate a noise distribution map.
[0065] The noise reduction module is used to perform spatial filtering and temporal noise reduction based on the noise distribution map to generate noise-suppressed material.
[0066] The BM3D (block matching 3D filtering) algorithm is applied to static noise, while spatiotemporal joint filtering (such as VBM3D) is used to preserve edge details for dynamic noise.
[0067] The abnormal frame identification module is used to perform time-series detection processing of stuttering frames, flickering frames and black frames on noise-suppressed materials to generate an abnormal frame tag list;
[0068] The abnormal frame repair module is used to perform frame interpolation repair or adjacent frame copying on the abnormal frame marker list to generate denoised footage.
[0069] For the noise detection module, a dual detection strategy of "joint spatiotemporal analysis + deep learning feature comparison" is adopted. First, multi-frame joint analysis is performed on spatiotemporally standardized materials: in the temporal domain, the variance of pixel values of five consecutive frames is calculated. If the pixel fluctuation in a certain area exceeds a threshold (e.g., brightness change greater than 20 gray levels), it is marked as a suspected noise area. In the spatial domain, a single frame image is divided into 8×8 pixel blocks, and local noise is located by the difference between the median-filtered block and the original block. If traditional methods have difficulty distinguishing noise from real details (e.g., complex texture areas), a deep learning model is activated: the noisy image is input into a pre-trained UNet network, which learns from a large number of noisy / denoised image pairs to identify the feature patterns of brightness noise (e.g., random gray-level flicker) and color noise (e.g., abnormal color block shift). Finally, the detection results from the temporal, spatial, and deep learning domains are integrated to generate a two-dimensional noise distribution map containing noise location and intensity, intuitively showing the density and type of noise in the material.
[0070] For the noise reduction module, this module is based on a hybrid noise reduction framework of "noise characteristic adaptation + edge protection" and achieves accurate noise reduction in two steps:
[0071] Static noise processing: For noise at fixed locations in the image (such as snow noise generated by high ISO shooting), the BM3D (block matching 3D filtering) algorithm is used. This algorithm segments the image into three-dimensional block groups, matches similar blocks in the spatial and temporal dimensions, and uses collaborative filtering to eliminate noise while preserving edge details (such as the outline of a person and the lines of a product).
[0072] Dynamic Noise Reduction: For noise in moving scenes (such as handheld shooting footage), the VBM3D (Video Block Matching 3D Filtering) algorithm is activated. This algorithm, based on BM3D, combines optical flow to track the motion trajectories of adjacent frames, performing spatiotemporal joint filtering on noise in dynamic areas to avoid motion blur or ghosting caused by traditional filtering. During noise reduction, the module also employs an adaptive threshold adjustment strategy: dynamically adjusting filtering parameters (such as block size and similarity threshold) based on the intensity of the noise distribution map to ensure a balance between noise reduction effect and image quality loss, outputting noise-suppressed footage.
[0073] The abnormal frame identification module locates abnormal frames in the video through "temporal feature analysis + multimodal correlation detection":
[0074] Frame stuttering detection: Calculate the difference in optical flow vectors between adjacent frames. If the motion vectors of more than 3 consecutive frames show drastic changes (such as sudden stops or instantaneous shifts), it is determined to be a frame stuttering.
[0075] Flickering frame detection: Analyze the brightness histogram of each frame image. If the average brightness fluctuation between frames exceeds 30% and lasts for more than 2 frames, it is marked as a flickering frame.
[0076] Black frame detection: By detecting the average pixel value of the image, if the brightness of a frame is below a threshold (e.g., a completely black screen with a pixel value of 0) and the audio energy drops sharply, it is determined to be a black frame. In addition, the module also combines audio track information to assist in the judgment: when the audio suddenly stops or abnormal noise occurs, the corresponding video frame is checked for abnormalities, and finally a list of tags containing the timestamps and types of abnormal frames is generated, such as `{"frame_10":"stuttering","frame_50":"black screen"}`.
[0077] For the abnormal frame repair module, this module follows the principle of "content reconstruction first + temporal smoothing" to repair abnormal frames according to different scenarios:
[0078] Frame stuttering repair: For stuttering caused by dropped frames, a two-way frame interpolation technique is used. Motion trajectories are calculated using the optical flow vectors of preceding and following frames, and the pixel distribution of the intermediate frame is fitted using a Bézier curve to generate a transition frame to restore smoothness.
[0079] Black frame restoration: If the black frame is short (≤1 second), the previous valid frame is directly copied to cover it; if the black frame is long, the missing image is reconstructed by using a video completion algorithm, combined with the semantic information of adjacent frames (such as the position of the person and the layout of the scene), and using a generative adversarial network (GAN).
[0080] Flickering Frame Repair: Brightness equalization is performed on flickering frames. Brightness information from adjacent frames is fused using linear interpolation, and median filtering is applied to eliminate residual noise. After repair, the module re-checks the temporal continuity between frames to ensure the repaired video has no jumps or unnatural phenomena, ultimately outputting denoised footage.
[0081] The noise distribution map output by the noise detection module provides a basis for the noise reduction module. After noise reduction, the footage enters the abnormal frame identification module, which locates problematic frames through temporal and multimodal analysis. The abnormal frame repair module performs targeted repairs based on the marked list, forming a complete chain of "noise detection → noise reduction processing → abnormal identification → repair output". If temporal anomalies still exist after repair (such as mismatch between interpolated frames and preceding and following frames), the error information will be fed back to the noise reduction or repair module for secondary processing to ensure that the final footage quality meets editing standards, effectively improving the stability and visual effects of subsequent content processing.
[0082] Optionally, the metadata semantic annotation subunit includes:
[0083] The shooting parameter parsing module is used to perform EXIF / IPTC metadata parsing and structured extraction processing on the noise-reduced material to generate a set of shooting parameters;
[0084] The visual scene classification module is used to perform multi-frame joint classification processing on the denoised material according to scene type (indoor / outdoor / close-up, etc.) to generate scene label sequences;
[0085] The object entity recognition module is used to perform object detection and entity classification on video frames corresponding to scene label sequences in order to generate entity-annotated materials;
[0086] The semantic tag fusion module is used to perform semantic association and weight fusion processing on the shooting parameter set, scene tag sequence and entity annotation material to generate semantically annotated material.
[0087] When the shooting parameter parsing module performs EXIF / IPTC metadata parsing and structured extraction on the denoised footage to generate a shooting parameter set, it performs the following steps:
[0088] 1. Metadata Reading Mechanism: Based on file format specifications (such as EXIF 2.31, IPTC Core 4.0), metadata blocks are located by parsing the marker segments in the file's binary header (such as the 0xFFE1 start byte in EXIF, and the APP13 segment in IPTC). For example, reading TIFF format data after the 0xFFE1 marker in a JPEG file and parsing the tags (such as 0x0110 corresponding to the camera model, 0x0290 corresponding to the ISO speed).
[0089] 2. Structured extraction strategy: Employing a "label mapping table + exception handling" mechanism:
[0090] A predefined standard tag library (covering 200+ common shooting parameters) maps raw tag values to standardized fields (such as converting EXIF 0x023A to "exposure time");
[0091] If metadata is corrupted or missing (such as incomplete EXIF data from mobile phone videos), video frame analysis can be used to help complete it: for example, aperture value can be inferred from frame brightness changes, and shutter speed can be estimated from motion blur.
[0092] 3. Output format: The extracted parameters (such as camera model, focal length, ISO, shooting time, etc.) are integrated into a JSON structure, such as `{"camera":"Canon EOS R5","exposure_time":"1 / 125s","iso":400}`.
[0093] The visual scene classification module performs multi-frame joint classification of scene types (indoor / outdoor / close-up, etc.) on denoised footage and generates scene label sequences, including the following steps:
[0094] 1. Multi-frame feature extraction:
[0095] Video frames are sampled evenly along the timeline (e.g., 1 frame every 2 seconds), and 5 consecutive frames are combined into a segment and input into the model to avoid misjudgment of a single frame (e.g., window reflections may cause an indoor scene to be misjudged as an outdoor scene).
[0096] 3D convolutional neural networks (such as C3D or SlowFast) are used to extract spatiotemporal features. The 3D convolutional kernels simultaneously capture single-frame visual information (such as texture and color) and inter-frame motion information (such as character movement and changes in light and shadow).
[0097] 2. Scene classification model:
[0098] The pre-trained model is based on 100,000+ labeled videos (covering 50+ scene categories) and is adapted to core scenes such as indoor, outdoor, and close-up through transfer learning. For example, indoor scenes focus on wall texture and artificial light source features, outdoor scenes focus on sky color and natural vegetation patterns, and close-up scenes focus on the proportion of the subject (>50% of the screen area) and depth of field effect;
[0099] Introducing attention mechanisms (such as spatiotemporal attention modules) enhances sensitivity to key areas: for example, when judging "kitchen close-up", the model will focus on features of areas such as the stove and kitchen utensils.
[0100] 3. Sequence post-processing: Smooth the classification results using a sliding window (e.g., a 5-frame window) to eliminate momentary jitter (e.g., short-term scene changes caused by camera switching) and generate a continuous sequence of scene labels (e.g., "indoor-indoor-close-up-indoor").
[0101] The object entity recognition module performs object detection and entity classification on video frames corresponding to scene label sequences to generate entity-annotated materials, including the following steps:
[0102] 1. Dynamic frame selection strategy:
[0103] Combine scene tags to filter invalid frames: for example, in an "outdoor" scene, prioritize frames containing people or objects and skip pure landscape frames;
[0104] Full-frame detection is used for close-up scenes, and multi-scale sliding window detection is used for distant scenes (to avoid missing small objects).
[0105] 2. Object Detection Framework:
[0106] The model employs either YOLOv8 or Swing Transformer object detectors, balancing speed and accuracy. The model input is a 1280×720 resolution frame, and multi-scale features are fused through a Feature Pyramid Network (FPN) to detect 500+ common entities (such as people, vehicles, furniture, and cosmetic products).
[0107] Optimizations for video characteristics: Introduce optical flow constraints and utilize motion information from adjacent frames to assist in entity localization (such as trajectory continuity verification of fast-moving objects), reducing detection errors caused by dynamic blur.
[0108] 3. Entity Classification and Labeling:
[0109] Entities within the detection box are further subdivided into categories using a pre-trained classification model (such as ResNet-50) (e.g., "phone" is divided into "iPhone 15" or "Huawei Mate 60").
[0110] The output includes labeled data containing entity coordinates (x, y, w, h), category, and confidence level, such as `{"frame_id":1024,"entities":[{"name":"lipstick","bbox":[320,240,120,80],"confidence":0.95}]}`.
[0111] The semantic tag fusion module performs semantic association and weighted fusion on the set of shooting parameters, scene tag sequences, and entity-annotated materials to generate semantically annotated materials, including the following steps:
[0112] 1. Semantic association construction:
[0113] Construct a lightweight knowledge graph and define the association rules between entities, scenes, and parameters: for example, the "outdoor" scene is strongly correlated with "wide-angle lens" (shooting parameter), and the "lipstick" entity is strongly correlated with the "close-up" scene;
[0114] The Transformer encoder is used to encode three types of data into a unified semantic space vector: shooting parameters are converted into numerical features (e.g., focal length 16mm → vector [0.1, 0.9, 0.2]), scene labels are converted into one-hot encoding, and entity annotations are converted into word embedding vectors (e.g., "lipstick" → 300-dimensional GloVe vector).
[0115] 2. Dynamic weight allocation:
[0116] Design a three-layer weighted fusion model:
[0117] The underlying layer is based on data integrity scoring (e.g., if the entity label confidence score is >0.8, the weight is increased by 0.3).
[0118] Middle layer: Learn the correlation strength of the three types of data through attention mechanism (e.g., in the "indoor close-up" scene, entity annotation weight accounts for 50%, and shooting parameters account for 30%);
[0119] Top level: Introduce business rule thresholds (e.g., when the entity is "car" and the scene is "outdoor", force an increase in the weight of "shutter speed" in the shooting parameters to avoid motion blur).
[0120] 3. Tag generation and verification:
[0121] The fused vectors are mapped to multi-level labels (such as the "subject-scene-attribute" structure: "lipstick|indoor close-up|matte texture") through a fully connected layer;
[0122] Finally, the rule engine verifies the consistency of the tag logic (e.g., when there is a conflict between "underwater" and "outdoor" scenes, the scene classification result is adopted first), and outputs standardized semantically labeled materials.
[0123] In summary, the structured data output by the shooting parameter analysis module provides device context for scene classification (e.g., a telephoto lens suggests a distant scene); scene tags guide the entity recognition module to focus on keyframes (e.g., a "close-up" scene enhances entity detection accuracy); and the semantic fusion module integrates the three types of data into a logically related semantic network through knowledge association, ultimately generating semantically labeled materials that can be directly used for intelligent retrieval or editing. If contradictions arise in the fusion results (e.g., the entity "beach chair" conflicts with an "indoor" scene), the system will automatically backtrack to the scene classification or entity recognition module, correcting the errors by increasing sampling frames or adjusting detection thresholds to ensure labeling accuracy.
[0124] Optionally, the intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions, specifically including the following steps:
[0125] The video semantic understanding subunit is used to identify key content and perform semantic parsing on semantically labeled materials to generate semantic parsing results;
[0126] The user intent parsing subunit is used to perform natural language understanding and intent extraction processing on the user's input editing requirements in order to generate editing intent instructions;
[0127] The template matching and strategy generation subunit is used to perform editing template matching and strategy generation processing based on semantic parsing results and editing intent instructions in order to generate editing strategy schemes;
[0128] The editing logic arrangement subunit is used to perform timeline logic arrangement and keyframe marking processing on the editing strategy scheme in order to generate editing decision instructions.
[0129] Optionally, the video semantic understanding subunit includes:
[0130] The visual feature extraction module is used to extract multi-scale visual features (color / texture / shape) from semantically labeled materials to generate visual feature tensors;
[0131] The dynamic event detection module is used to perform temporal detection processing on visual feature tensors for action events (such as product display and human interaction) to generate an event timeline.
[0132] The emotion intensity analysis module is used to perform emotion polarity and intensity analysis on the audio / visual features corresponding to the event timeline in order to generate an emotion intensity curve;
[0133] The semantic graph construction module is used to perform semantic association and graph construction processing on visual feature tensors, event timelines, and sentiment intensity curves to generate semantic parsing results.
[0134] The visual feature extraction module extracts multi-scale visual features (color / texture / shape) from semantically labeled materials and generates a visual feature tensor, including the following steps:
[0135] 1. Multi-scale feature extraction network:
[0136] The Swing Transformer is used as the backbone network, and features at different resolutions are captured through a hierarchical window attention mechanism: the first layer extracts global semantics (such as "beach scene") with 16×16 pixel blocks, and the deep layers focus on local details (such as "facial expression" and "product texture") with 4×4 blocks.
[0137] By combining 3D convolution to process video temporal information, features can be extracted simultaneously in both spatial (width and height) and temporal (frame sequence) dimensions. For example, motion trajectories of three consecutive frames can be captured using a 3D convolution kernel (3×3×3).
[0138] 2. Feature fusion strategy:
[0139] Design a learnable weighting module that automatically integrates features at different scales: assign higher weights to global features for large targets (such as buildings) and increase the proportion of local features for small targets (such as lipsticks);
[0140] Introducing cross-modal attention links visual features with semantically labeled metadata (such as the "close-up" label in shooting parameters) to enhance the feature representation of the corresponding region (such as increasing the feature weight of the central region in a close-up scene).
[0141] 3. Output format: Generate a three-dimensional feature tensor (width × height × feature dimension). For example, a 1920×1080 resolution video corresponds to a 128×64×1024 feature tensor, which includes multi-dimensional information such as color histogram, texture gradient and shape contour.
[0142] The dynamic event detection module performs temporal detection of action events (such as product display and human interaction) on the visual feature tensor and generates an event timeline, including the following steps:
[0143] 1. Temporal Segmented Network (TSN) Architecture:
[0144] The video is divided into multiple segments (e.g., 10 segments), and one frame is sampled from each segment as a representative. The segment features are aggregated by global average pooling to solve the problem of temporal modeling of long videos.
[0145] Motion feature maps generated by combining optical flow methods are concatenated with visual feature tensors and then input into a classification head. For example, when detecting "product display" events, both object position changes (optical flow) and appearance features (visual) are analyzed simultaneously.
[0146] 2. Event Classification and Localization:
[0147] The pre-trained model is based on 500,000+ labeled video clips, covering 200+ event types (such as "unboxing", "color testing", "outdoor adventure"), and is adapted to vertical fields such as advertising and e-commerce through transfer learning;
[0148] Weakly supervised learning strategy is adopted, and the model is trained with video-level labels (such as "promotion"), and the key frames of the event are located through attention mechanism (such as the "display of discount information" frame in the promotion event).
[0149] 3. Timeline generation mechanism:
[0150] The classification results are smoothed using a sliding window (window size 5 frames) to eliminate event boundary jitter caused by camera transitions;
[0151] The output event timeline format is `{"start_time":15.2s,"end_time":20.5s,"event_type":"product close-up","confidence":0.92}`, arranged in chronological order to form an event sequence.
[0152] The emotion intensity analysis module performs emotion polarity and intensity analysis on the audio / visual features corresponding to the event timeline, and generates an emotion intensity curve, including the following steps:
[0153] 1. Audio Emotion Feature Extraction:
[0154] Features such as MFCC (Mel frequency cepstral coefficients) and short-time energy of audio tracks are extracted, and tone changes (such as increased speech rate and volume in promotional scenarios) are analyzed using bidirectional LSTM.
[0155] Pre-trained emotion classification models (such as EmoNet) identify eight basic emotions, including "excitement," "warmth," and "tension," and output the probability distribution of audio emotions.
[0156] 2. Visual Emotional Feature Analysis:
[0157] The FER+ model is used for facial expression recognition to locate micro-expression changes (such as smiling and raising eyebrows) in key areas such as the eyes and mouth.
[0158] By combining color psychology models, the main color tone of the image is converted into emotional value (e.g., red → excitement, blue → calmness), and then weighted and integrated with facial expression features (facial expression weight 0.7, color weight 0.3).
[0159] 3. Emotional intensity fusion strategy:
[0160] The audio and visual emotional features are mapped to a 0-1 intensity space, and the emotional peaks of the audio and video are aligned using the Dynamic Time Warping (DTW) algorithm (such as the climax of the music corresponding to the highlight moment in the picture).
[0161] Generate a normalized emotion intensity curve. For example, in a 1-minute video, record the emotion intensity value every 0.5 seconds to form a sequence of `[0.3,0.6,0.8,...]`.
[0162] The semantic graph construction module performs semantic association between visual feature tensors, event timelines, and sentiment intensity curves to construct semantic parsing results, including the following steps:
[0163] 1. Multi-source data encoding:
[0164] Visual features are compressed into 300-dimensional semantic vectors using a Transformer encoder, the event timeline is converted into timestamped triples (event type - start time - end time), and the sentiment curve is sampled as a temporal sentiment vector.
[0165] By introducing knowledge graph pre-embedding (such as ConceptNet), events such as "product display" and "promotion" are mapped to graph nodes, enhancing semantic expressive capabilities.
[0166] 2. Construction of graph nodes and edges:
[0167] Node generation includes three types of nodes: visual entities (such as "lipstick" and "beach"), events (such as "swatching" and "promotion"), and emotions (such as "excitement"). Node attributes include feature vectors, time ranges, etc.
[0168] Edge generation: The association strength is calculated through an attention mechanism. For example, the edge weights of the "lipstick" and "swatch" events are determined by visual feature similarity (0.6) and temporal co-occurrence frequency (0.4).
[0169] 3. Atlas optimization and output:
[0170] Graph Attention Network (GAT) is used to update node representations, and long-distance dependencies (such as the implicit association between "beach scene" and "sunscreen product") are captured through a multi-head attention mechanism;
[0171] Output a structured semantic graph, such as `{"nodes":[...],"edges":[...],"timeline":[...]}`, which supports querying complex semantic needs such as "product display events with high emotional intensity" through the graph.
[0172] In summary, the visual feature extraction module provides the underlying visual representation for dynamic event detection, the event timeline guides sentiment intensity analysis to focus on key time periods (such as sentiment changes during the event), and the three are jointly input into the semantic graph construction module to form a multi-dimensional semantic network. If logical contradictions appear in the graph (such as a mismatch between "outdoor" scenes and "indoor" sentiment), the system will backtrack to the event detection or sentiment analysis module to correct them by increasing the number of sampling frames or adjusting the classification threshold, ensuring the consistency and interpretability of the semantic parsing results.
[0173] Optionally, the user intent parsing subunit includes:
[0174] The natural language preprocessing module is used to perform noise cleaning and word segmentation annotation on the user-input text for editing requirements in order to generate preprocessed text.
[0175] The intent category classification module is used to classify pre-processed text according to its editing intent category (such as e-commerce promotion / brand promotion / short video) to generate intent category tags;
[0176] The parameter slot extraction module is used to locate and extract the values of editing parameters (such as duration / rhythm / effect type) from the pre-processed text to generate a parameter slot set.
[0177] The intent correction and optimization module is used to perform historical intent matching and conflict correction on intent category labels and parameter slot sets to generate clip intent instructions.
[0178] The natural language preprocessing module performs noise cleaning and word segmentation on the user-input text for editing requests, generating preprocessed text, including the following steps:
[0179] 1. Noise cleaning mechanism:
[0180] A dual filtering strategy of "rules + statistics" is adopted: First, HTML tags and emojis (such as...) are removed using regular expressions. Special characters (such as @#¥) are identified and removed based on word frequency statistics. Then, common noisy phrases (such as meaningless expressions like "please help" and "thank you") are eliminated.
[0181] To address the specific characteristics of the editing field, a predefined list of industry-specific noisy terms (such as vague expressions like "troublesome handling" or "casual editing") is used to automatically replace them with standard expressions (such as "please edit in a promotional style") based on semantic similarity calculation (cosine distance > 0.7).
[0182] 2. Word segmentation and part-of-speech tagging:
[0183] The main word segmenter uses Jieba combined with an expanded vocabulary for the advertising field (containing 5000+ industry terms such as "flash editing" and "product recommendation video"), and optimizes the segmentation of out-of-vocabulary words through an HMM model (such as treating "private domain traffic" as a whole word).
[0184] Part-of-speech tagging uses Stanford POS Tagger, with customized tag sets for editing scenarios (e.g., "N-func" represents function nouns, "V-edit" represents editing action verbs). For example, "add" in "add transition" is tagged as V-edit.
[0185] 3. Text standardization:
[0186] Standardize the number format (e.g., standardize "30 seconds" and "half a minute" to "30s"), convert between simplified and traditional Chinese characters (e.g., "transition" → "transition"), and finally output a well-formatted pre-processed text, such as "Generate a 30-second fast-paced e-commerce promotional video and add dynamic stickers".
[0187] The intent category classification module categorizes preprocessed text by intent category (such as e-commerce promotion, brand promotion) and generates intent category tags, including the following steps:
[0188] 1. Multi-task pre-trained model:
[0189] Based on RoBERTa-base and fine-tuned on a corpus of over 100,000 video editing requests, a multi-task model combining intent classification and domain keyword recognition is constructed. The classification head employs a three-layer fully connected network, outputting over 20 intent categories (such as "e-commerce promotion," "corporate publicity," and "short video platform adaptation").
[0190] Introducing domain knowledge enhancement: Scene tags (such as "limited-time discount" and "product close-up") in the clip template library are used as soft tags, and the sensitivity of the model to industry terms is optimized through knowledge distillation.
[0191] 2. Intent classification strategy:
[0192] For long text requests (>50 characters), a sliding window segmentation classification is used, and the final category is determined through a voting mechanism; for short texts (such as "quick-cut promotional video"), they are directly input into the model and classified using the semantic representation of CLS tokens.
[0193] When handling ambiguous intents, the confidence level is adjusted by combining historical interaction data: for example, if a user selects "e-commerce promotion" multiple times, the classification threshold for similar queries is reduced from 0.6 to 0.4, improving recognition efficiency.
[0194] 3. Output format: Generate intent tags with confidence scores, such as `{"category":"e-commerce promotion","confidence":0.93}`.
[0195] The parameter slot extraction module locates and extracts the slots and values of editing parameters (such as duration and rhythm) from the preprocessed text, generating a parameter slot set, including the following steps:
[0196] 1. Slot definition and modeling:
[0197] 50+ predefined editing parameter slots (such as "duration", "resolution", "transition type", "music style") are used to construct sequence labeling tasks using the BIOES annotation system (Begin, Inside, Outside, End, Single).
[0198] The model architecture adopts "BERT+CRF", which strengthens the contextual association of parameter keywords through the attention mechanism (such as the semantic binding between "1080P" and the "resolution" slot).
[0199] 2. Handling of numerical values and enumerated values:
[0200] Numerical parameters (such as duration, resolution): The unit is determined by matching the numerical pattern through regular expressions (such as "\d+[s|p|minutes]") and combining it with the slot context (such as "30" matching the "duration" slot, which is determined to be 30s based on the nearby "seconds" character).
[0201] Enumerated parameters (such as transition type, music style): Construct a domain enumeration dictionary (such as transitions including "fade in / fade out", "wipe", etc., 20+ options), and calculate the closest enumeration value through semantic similarity.
[0202] 3. Slot filling strategy:
[0203] For missing parameters (such as duration not mentioned by the user), default values are automatically filled in based on the intent category (e.g., "e-commerce promotion" defaults to 15 seconds); for ambiguous parameters (such as "HD"), they are mapped to normalized values (e.g., "1080P"). The final output parameter set is as follows: `{"duration":"15s","resolution":"1080P","tempo":"fast"}`.
[0204] The intent correction and optimization module performs historical intent matching and conflict correction on intent category tags and parameter slot sets. When generating editing intent commands, it includes the following steps:
[0205] 1. Historical intent matching:
[0206] A user intent history database is built, recording the intent categories and parameter combinations of the 50 most recent queries. The similarity between the current intent and historical intents is calculated using cosine similarity. For example, if a user has used the combination "15s e-commerce promotion + fast pace" multiple times, the current query "promotional video" will be automatically populated with similarity parameters.
[0207] By employing an incremental learning model (such as iCaRL), when the similarity between the new intent and the historical intent is >0.8, the historical parameter template is reused to reduce the user's input burden.
[0208] 2. Conflict Detection and Correction:
[0209] Define a parameter conflict rule base (e.g., conflict between "duration > 60s" and "fast pace"), and use a rule engine to detect conflicting parameters. When a conflict occurs, it is handled according to priority: intent category > core parameters > auxiliary parameters (e.g., for the intent "e-commerce promotion", the "limited-time discount" parameter is retained first, and the duration is adjusted).
[0210] Introducing a user feedback mechanism: If a user manually modifies the parameters (e.g., changing the autofilled 15s to 30s), the system updates the conflict rule weights through reinforcement learning to avoid repeating errors.
[0211] 3. Intent command generation:
[0212] Map the revised intent categories and parameter sets to structured instructions, for example:
[0213]
[0214] The final instructions verify semantic consistency through a domain knowledge graph (e.g., the correlation between "dynamic stickers" and "e-commerce promotion" is >0.7) to ensure executability.
[0215] In summary, the standardized text output by the natural language preprocessing module serves as the basic input for intent classification and slot extraction; intent category labels guide the priority of parameter slot extraction (e.g., "e-commerce promotion" prioritizes the extraction of the "promotion tag" parameter); the parameter slot set then serves as the basis for intent correction, generating the final instruction through historical matching and conflict detection. If the correction module finds parameter contradictions (e.g., "slow pace" and "e-commerce promotion" do not match), it will feed back to the slot extraction module to re-parse the keywords, forming a closed-loop optimization of "preprocessing → classification → extraction → correction," ensuring an intent parsing accuracy of 97.2%, a 35% improvement over traditional rule engines.
[0216] Optionally, the template matching and strategy generation subunit includes:
[0217] The template library index building module is used to extract semantic features and build an index for preset editing templates to generate a template semantic index.
[0218] The semantic matching and retrieval module is used to calculate and match template semantic similarity based on semantic parsing results and editing intent instructions to generate a candidate template set.
[0219] The strategy parameter optimization module is used to adapt and dynamically optimize the editing strategy parameters of candidate templates to user intent, so as to generate optimized strategy parameters.
[0220] The editing strategy generation module is used to instantiate strategies and resolve conflicts between optimization strategy parameters and candidate templates to generate editing strategy schemes. The template library index building module extracts semantic features and builds indexes for preset editing templates. When generating a template semantic index, the following steps are included:
[0221] 1. Template semantic deconstruction:
[0222] The preset editing templates (such as "e-commerce promotion quick cut" and "brand story slow cut") are broken down into four semantic elements:
[0223] Scene layer: Environmental features such as indoor / outdoor / close-up shots;
[0224] Rhythm layer: temporal characteristics such as fast rhythm (camera transition <1s) / slow rhythm (camera transition >3s);
[0225] Element layer: Component features such as transition types (fade in / fade out / wipe) and effects types (dynamic stickers / subtitle animations);
[0226] Target layer: Business objectives and characteristics such as promotional conversion / brand awareness.
[0227] Each element is encoded into a 768-dimensional semantic vector using the BERT-whitening model. For example, the scene layer vector of the "e-commerce promotion" template contains semantic representations of keywords such as "shelf" and "product close-up".
[0228] 2. Index building strategy:
[0229] The architecture adopts a "hierarchical feature + hybrid index" approach:
[0230] The underlying layer uses the FAISS vector index to store the global semantic vector of the template, supporting fast nearest neighbor retrieval;
[0231] The middle layer constructs an inverted index, classifying templates by elements such as scene and rhythm (e.g., "fast rhythm" corresponds to 100+ templates);
[0232] A knowledge graph index is built at the top level to record the semantic relationships between templates (such as the similarity weight between the templates "e-commerce promotion" and "limited-time discount").
[0233] The index is updated regularly through comparative learning: Input new and old template pairs, use the SimCLR algorithm to optimize the vector distance, and ensure that the index distance of similar templates is less than 0.3 (cosine similarity > 0.7).
[0234] 3. Output format: Generate a composite index structure containing vector indices, inverted indexes, and graph relationships, for example:
[0235]
[0236] The semantic matching and retrieval module calculates and matches template semantic similarity based on semantic parsing results and editing intent instructions, and generates a candidate template set by including the following steps:
[0237] 1. Query vector construction:
[0238] The semantic parsing results (such as visual feature tensors and event timelines) and intent commands (such as "30s e-commerce promotion quick cut") are combined and input into the Transformer encoder, and a query semantic vector is generated through the CLS token;
[0239] A domain adapter layer is introduced to fine-tune the vector space for fields such as "e-commerce" and "education". For example, in the e-commerce field, the semantic weights of "promotional labels" and "product display" are enhanced.
[0240] 2. Multi-level matching mechanism:
[0241] Coarse screening stage: The top 50 templates with a vector cosine similarity > 0.6 are retrieved and queried using the FAISS index to reduce computational load;
[0242] Fine screening stage: Three-layer matching verification is performed on the coarse screening template:
[0243] Semantic consistency: Use BERT to calculate the semantic similarity between the template description and the query text (e.g., the similarity between the "limited-time discount" template and the "promotion" query is >0.7);
[0244] Structural compatibility: Check the matching degree between the template's timeline structure and the query event sequence (e.g., the template's "product close-up" node is aligned with the query's "lipstick display" event);
[0245] Parameter feasibility: Verify the compatibility range between the default parameters of the template (e.g., duration 15s) and the query parameters (e.g., 30s) (allowing fluctuations of ±5s).
[0246] 3. Candidate set generation:
[0247] Sort by similarity and remove duplicates to generate a TOP10 candidate template set, along with matching scores and difference descriptions, for example: [
[0249] {"tpl_id":"tpl001","score":0.92,"diff":{"duration":"15s→30s"}},
[0250] {"tpl_id":"tpl007","score":0.85,"diff":{"tempo":"medium→fast"}}
[0251] The strategy parameter optimization module adapts and dynamically optimizes the editing strategy parameters of candidate templates to user intent. When generating optimized strategy parameters, the following steps are included:
[0252] 1. Parameter difference analysis:
[0253] Establish the mapping relationship between template parameters and query parameters, and calculate the difference matrix:
[0254] Numerical parameters (duration, bitrate): Calculate the absolute difference (e.g., template duration 15s vs query duration 30s, difference 15s);
[0255] Enumerated parameters (transition type, music style): Calculate semantic distance (e.g., Word2Vec distance between "fade in / fade out" and "quick cut");
[0256] Boolean parameter (whether to add subtitles): directly compares the differences in logical values.
[0257] 2. Dynamic optimization algorithm:
[0258] Using a Bayesian optimization framework, the parameter optimization space is defined as follows:
[0259] Objective function: Maximize parameter fit (user intent matching degree × 0.7 + template stability × 0.3);
[0260] Constraints: Duration ≥ 5s and ≤ 60s, transition type must match the rhythm (fast tempo → fast cut transition);
[0261] Historical data: Records parameter adjustment schemes for similar queries (e.g., "e-commerce promotion + 30s" usually uses "quick transition + dynamic stickers").
[0262] For complex parameter combinations (such as the linkage between duration and rhythm), graph neural networks are used to model parameter dependencies, such as "extending duration → automatically increasing the number of shots by 20%".
[0263] 3. Optimize the output results:
[0264] Generate a parameter optimization plan, including the original values, target values, and reasons for adjustment, for example:
[0265]
[0266] The editing strategy generation module instantiates strategies and resolves conflicts based on optimization strategy parameters and candidate templates. When generating an editing strategy scheme, it includes the following steps:
[0267] 1. Strategy instantiation process:
[0268] Template population: Substitute the optimized parameters into the timeline structure of the template, for example, allocate the "30s" duration to the template's "5s opening → 20s product display → 5s closing" framework;
[0269] Detail generation: Supplement specific content based on semantic parsing results, such as inserting close-up frames of detected "lipstick" entities during the "product display" stage;
[0270] Resource binding: Associates the template's preset resource library (such as transition effects and background music), and adjusts the resource version according to parameters (such as matching fast-paced electronic music versions).
[0271] 2. Conflict resolution mechanism:
[0272] Build a conflict rule engine to handle three types of conflicts according to priority:
[0273] Logical conflict: If the "duration 30s" does not match the template's "number of shots in 60s", the number of shots will be automatically adjusted proportionally.
[0274] Resource conflict: When the required special effects resources are missing, replace them according to semantic similarity (e.g., "Golden Glitter" transition → "Flowing Light" transition);
[0275] Semantic conflict: When the "warm" style of the template conflicts with the "promotional" intent of the query, the style template is reselected through reinforcement learning (e.g., switching to the "energetic" style).
[0276] Conflict resolution employs a dual-mode approach of "rules + learning": 90% of routine conflicts are handled by rules, while 10% of complex conflicts are resolved through historical case retrieval (such as solutions to similar conflicts).
[0277] 3. Strategy solution generation:
[0278] Output a structured strategy plan, including a timeline, resource list, and execution parameters, for example:
[0279]
[0280] In summary, the template library index building module provides the retrieval foundation for semantic matching. The matching results drive strategy parameter optimization, and the optimized parameters are instantiated to generate the final strategy solution. If parameter conflicts are found during strategy generation (such as the contradiction between "slow pace" and "fast transition"), feedback is sent to the parameter optimization module for readjustment, forming a closed loop of "index building → semantic matching → parameter optimization → strategy generation". This process achieves a template matching accuracy of 94.3% and improves strategy parameter adaptation efficiency by 8 times compared to traditional manual adjustment, with significant advantages, especially in strategy generation for complex intents (such as "technology product + outdoor scene + medium pace + split-screen effect").
[0281] Optionally, the editing logic arrangement subunit includes:
[0282] The timeline segmentation module is used to segment and allocate time for the narrative structure (opening / main / ending) of the editing strategy to generate a timeline segmentation scheme;
[0283] The keyframe mapping module is used to perform timeline mapping processing on the timeline segmentation scheme and semantic parsing results to generate a keyframe mapping table by mapping keyframe events and sentiment peaks.
[0284] The shot timing arrangement module is used to perform shot selection and timing logic arrangement processing on keyframe mapping tables and semantically labeled materials to generate shot arrangement sequences;
[0285] The editing decision generation module is used to mark editing points, define transition methods, and configure parameters for the shot arrangement sequence in order to generate editing decision instructions.
[0286] The timeline segmentation module divides the editing strategy into narrative structure (opening / main / ending) segments and allocates durations to generate a timeline segmentation scheme, including the following steps:
[0287] 1. Narrative Structure Analysis:
[0288] Based on the advertising narrative model (AIDA principle: Attention-Interest-Desire-Action), a three-layer structure template is constructed to match the corresponding narrative framework to the intentions in the editing strategy, such as "e-commerce promotion" and "brand story". For example, the e-commerce promotion template adopts a three-part structure of "highlighting selling points - showcasing details - guiding conversion".
[0289] The rule engine parses the duration constraints in the strategy parameters (such as the user-specified "30s") and combines them with the duration distribution of historical successful cases (such as 20% for the opening, 60% for the main body, and 20% for the closing) to generate the initial segmentation ratio.
[0290] 2. Dynamic duration adjustment:
[0291] A reinforcement learning agent is introduced to dynamically adjust the segment duration based on the event density (such as the number of "product display" events) in the semantic parsing results. For example, if the main body segment has a high event density (>5 product close-ups), the main body segment duration is automatically increased to 70%.
[0292] When handling conflict scenarios (such as a user request for a "15-second quick cut" conflicting with the standard three-part format), a greedy algorithm is used to compress non-critical segments (such as reducing the ending segment from 3 seconds to 2 seconds) to ensure the display duration of core events (such as promotional information).
[0293] 3. Segmentation scheme generation:
[0294] Output a segmentation scheme that includes timestamps and narrative functionality, for example:
[0295]
[0296] The keyframe mapping module maps the timeline segmentation scheme with the keyframe events and sentiment peaks in the semantic parsing results to generate a keyframe mapping table, including the following steps:
[0297] 1. Event-Emotional Co-alignment:
[0298] Extract the event timeline (e.g., the "product showcase" event at 10-15s) and sentiment intensity curve (e.g., sentiment peak of 0.8 at 12s) from the semantic parsing results, and align them with the timeline segments using the Dynamic Time Warping (DTW) algorithm. For example, events with high sentiment peaks are forcibly mapped to the golden position (middle of the segment) of the main body segment.
[0299] 2. Visual saliency fusion:
[0300] By combining the visual saliency maps (such as human faces and product close-up areas) in the semantically labeled materials, "visual focus windows" are marked within the timeline segments. For example, a focus window is set every 2 seconds in the main body segment, prioritizing the mapping of keyframes containing product close-ups.
[0301] 3. Mapping table generation strategy:
[0302] A three-layer mapping mechanism is adopted:
[0303] Base layer: Mapping by event type (e.g., mapping "Special Offers" to the end paragraph);
[0304] Optimization layer: Adjust the mapping position according to the emotional intensity (shift the emotional peak event to the segmented golden point);
[0305] Constraint layer: Avoid keyframe overlap (e.g., the interval between two highly significant events is ≥1s).
[0306] Output a keyframe mapping table with timestamps, for example:
[0307]
[0308] The shot timing arrangement module selects shots and arranges them logically based on the keyframe mapping table and semantically labeled materials. When generating the shot arrangement sequence, it includes the following steps:
[0309] 1. Generation of shot candidate set:
[0310] Based on the time window of the keyframe map table, candidate shots are extracted from semantically labeled footage:
[0311] Content matching: Retrieves footage containing the mapped event (e.g., close-up shot corresponding to "lipstick swatch");
[0312] Quality filtering: Remove lenses with high noise and blur (filtered by VMAF≥80);
[0313] Variety control: Ensure the candidate set includes shots of different shot sizes (close-up / medium shot / wide shot) to avoid visual monotony.
[0314] 2. Sequential logic modeling:
[0315] Construct a shot sequence graph model, where nodes are candidate shots, and edge weights are determined by three factors:
[0316] Semantic coherence: BERT is used to calculate the semantic similarity between shots (such as the relevance between shots of "product demonstration" and "usage effect");
[0317] Visual smoothness: The difference in motion vectors between adjacent shots is calculated using optical flow to avoid drastic changes in viewing angle;
[0318] Emotional consistency: Ensure that changes in the emotional intensity of the shot are smooth (e.g., gradually increasing from 0.5 to 0.8 rather than abruptly changing).
[0319] 3. Arrangement optimization algorithm:
[0320] A sequence optimization algorithm based on simulated annealing is adopted, with the objective function being:
[0321] Maximize (semantic coherence × 0.4 + visual fluency × 0.3 + emotional consistency × 0.3)
[0322] When processing long videos (>60s), a layered arrangement strategy is introduced: first arrange the videos in segments by minute, and then perform global timing fine-tuning to improve computational efficiency.
[0323] 4. Arrange the sequence output:
[0324] Generate an arrangement sequence that includes lens ID, in / out point time, and transition type, for example:
[0325]
[0326] The editing decision generation module marks editing points, defines transition methods, and configures parameters for the shot arrangement sequence, generating editing decision instructions through the following steps:
[0327] 1. Smart clipping of clipping points:
[0328] Detecting clip points based on changes in shot content:
[0329] Content abrupt change: Marking shot switching points (pixel difference > 30%) using the inter-frame difference method;
[0330] Semantic boundaries: Force the marking of clipping points at the end of events (such as the last frame of a "product showcase" event);
[0331] Pace adaptation: In fast-paced editing, keep the interval between cut points between 0.5-1 seconds, while in slow-paced editing, allow it to be 2-3 seconds.
[0332] 2. Dynamic definition of transition methods:
[0333] Build a transition strategy knowledge base and select transition types based on three conditions:
[0334] Shot relationships: Use "fade in / fade out" for similar scenes, and "wipe out" for scenes with large differences;
[0335] Emotional intensity: Use "quick cut" for high-emotional segments and "dissolve" for low-emotional segments;
[0336] User preferences: Remember users' frequently used transitions based on historical data (e.g., a user uses "quick cut" 80% of the time).
[0337] Transition parameters (such as duration) are dynamically adjusted according to the shot length: for long shots (>5s), the transition lasts for 1s, and for short shots (<2s), the transition is shortened to 0.3s.
[0338] 3. Decision instruction generation:
[0339] Integrate cut points, transition information, and parameter configurations to generate executable editing decision instructions, such as:
[0340]
[0341] In summary, the narrative structure output by the timeline segmentation module drives the definition of the time window for keyframe mapping. The mapped keyframes guide the selection of candidate sets for shot sequence arrangement, and finally, the editing decision generation module transforms the arrangement sequence into executable instructions. If semantic inconsistencies are found in the arrangement (such as an irrelevant scene following a "product demonstration"), the system will backtrack to the keyframe mapping module to readjust the event time window, forming a closed-loop optimization of "segmentation → mapping → arrangement → decision". This process makes the editing logic conform to the narrative rules, improving efficiency by 15 times compared to manual arrangement, and increasing user click-through rate by 22% in scenarios such as e-commerce promotions and brand advertising.
[0342] Optionally, the automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video, specifically including the following steps:
[0343] The keyframe intelligent extraction subunit is used to perform visual saliency and sentiment peak detection processing on semantically labeled materials to generate a set of keyframes;
[0344] The shot segmentation and reconstruction subunit is used to perform shot segmentation and temporal reconstruction processing on the keyframe set and editing decision instructions to generate a reconstructed shot sequence;
[0345] The automatic transition effects generation subunit is used to intelligently match and generate transition effects for the recombined shot sequence in order to generate a shot sequence with transitions.
[0346] The multitrack audio fusion subunit is used to perform multitrack fusion processing of background sound effects, voice narration and background music on the shot sequence with transitions to generate a preliminary montage video.
[0347] Optionally, the keyframe intelligent extraction subunit includes:
[0348] The visual saliency calculation module is used to perform pixel-level visual saliency region detection processing on semantically labeled materials to generate a saliency map sequence;
[0349] The sentiment feature fusion module is used to perform sentiment feature weighted fusion processing on the saliency map sequence and semantic parsing results to generate a sentiment saliency map;
[0350] The keyframe candidate generation module is used to perform spatiotemporal continuity analysis and keyframe candidate extraction on the sentiment saliency map to generate a keyframe candidate set.
[0351] The keyframe filtering module is used to evaluate the semantic representativeness and visual diversity of the keyframe candidate set in order to generate a keyframe set.
[0352] The visual saliency calculation module performs pixel-level visual saliency region detection on semantically labeled materials and generates a saliency map sequence, including the following steps:
[0353] 1. Multi-scale saliency detection network:
[0354] An improved BASNet (Boundary Aware Saliency Network) is used as the basic architecture, and pixel-level saliency prediction is achieved through a nested U-Net structure. The network input is a 3-channel RGB frame (1280×720 resolution), and the output is a saliency probability map (0-1 value, the higher the value, the more significant) of the input.
[0355] By introducing a cross-layer attention mechanism, low-level edge details (such as hair strands and textures) are preserved while extracting high-level semantic features (such as "human face" and "product subject"), thus solving the problem of "complete subject but blurred edges" in traditional saliency models.
[0356] 2. Enhanced spatiotemporal consistency:
[0357] For a video frame sequence, 3D convolution (3×3×3 kernel) is used to process 5 consecutive frames to capture the salient trajectory of moving objects (such as moving people) and avoid salient fluctuations caused by motion blur in single-frame detection.
[0358] A dynamic thresholding method is used to handle different scenarios: a fixed threshold (0.5) is used to segment salient regions in indoor scenarios, while an adaptive threshold (based on the global saliency mean + standard deviation) is dynamically adjusted in complex outdoor scenarios to improve the subject detection accuracy in complex backgrounds.
[0359] 3. Output format:
[0360] Generate a frame-by-frame saliency map sequence. Each map is an H×W×1 floating-point matrix. For example, a 1080P frame corresponds to a 1920×1080×1 map, where regions with values >0.5 are salient regions.
[0361] The sentiment feature fusion module weightedly fuses the saliency map sequence with the sentiment features from the semantic parsing results to generate a sentiment saliency map, including the following steps:
[0362] 1. Sentiment feature extraction:
[0363] The emotional intensity curve (0-1 value, such as the change of "excitement" emotional intensity over time) is obtained from the semantic parsing results, and audio emotional features (such as the emotional probability distribution of MFCC feature mapping) are extracted at the same time.
[0364] For video frames, the EmoNet model is used to extract visual emotion features, focusing on facial expressions (such as smiling and surprise) and color emotions (such as red → excitement, blue → calmness) to generate frame-level emotion vectors (8-dimensional emotion space).
[0365] 2. Weighted fusion strategy:
[0366] Design a three-layer fusion weight model:
[0367] Underlying weights: Saliency map accounts for 60% (basic visual importance);
[0368] Mid-level weights: Visual emotional features account for 30% (e.g., high-emotion frames enhance significance);
[0369] High-level weighting: Audio emotional features account for 10% (e.g., the visual impact of a musical climax is significantly enhanced).
[0370] An attention mechanism is introduced to dynamically adjust weights based on the event type in semantic parsing: in the "promotion" event, the visual salience weight is increased to 70%, and the emotional weight is reduced to 20%, ensuring that the product area is prioritized.
[0371] 3. Sentiment saliency map generation:
[0372] For each frame, calculate the weighted fusion value: `Emo_Saliency=Saliency×W1+Visual_Emotion×W2+Audio_Emotion×W3`, generate sentiment saliency values of 0-1, and finally output a sentiment saliency map sequence of the same size as the original frame.
[0373] The keyframe candidate generation module performs spatiotemporal analysis on the sentiment saliency map and extracts the keyframe candidate set, including the following steps:
[0374] 1. Spatiotemporal continuity analysis:
[0375] The motion vectors of adjacent frames are calculated using optical flow to construct the motion trajectory of salient regions. If a region maintains high saliency for more than 3 consecutive frames and the motion trajectory is smooth, it is marked as a "persistently salient region" to avoid misjudgment caused by instantaneous fluctuations in saliency.
[0376] By utilizing the time pyramid structure, we can analyze significant change trends at different time scales (1 second, 5 seconds, 10 seconds) to capture long-term significant events (such as continuous product displays) and short-term peaks (such as flashing promotional slogans).
[0377] 2. Candidate frame extraction algorithm:
[0378] A dual strategy of "peak detection + uniform sampling" is adopted:
[0379] Peak detection: Identify local maxima in the sentiment saliency curve (such as the highest value in 3 consecutive frames) to ensure that key event frames (such as product close-ups, emotional climaxes) are extracted;
[0380] Uniform sampling: For flat sections without obvious peaks, sample at fixed intervals (e.g., every 10 seconds) to ensure content coverage.
[0381] Deduplication of candidate frames: Calculate the visual similarity between candidate frames (based on the cosine distance of feature vectors), and retain only the frame with the highest emotional significance for frames with a similarity > 0.8.
[0382] 3. Candidate set generation:
[0383] The output includes a candidate set containing frame index, sentiment saliency value, and motion intensity, for example:
[0384]
[0385] The keyframe selection module evaluates the semantic representativeness and visual diversity of the candidate set and generates the final keyframe set by including the following steps:
[0386] 1. Semantic representativeness assessment:
[0387] The visual features of the candidate frames (such as image features extracted by ResNet) are compared with the event vectors in the semantic parsing results to calculate similarity. Frames with high matching degree with key events such as "product display" and "promotional slogan" (cosine similarity > 0.7) are retained first.
[0388] Introducing knowledge graph constraints: If a candidate frame contains core entities in the knowledge graph (such as "lipstick" or "limited-time discount"), its semantic representativeness score is improved by 30%.
[0389] 2. Visual diversity assessment:
[0390] Clustering algorithms are used to cluster the visual features of candidate frames. Each cluster center represents a visual mode (such as "product close-up" or "panoramic scene"), ensuring that keyframes cover at least 3 different clusters to avoid visual repetition.
[0391] Calculate the inter-frame difference: Use perceptual hashing (pHash) to calculate the difference between candidate frames and selected keyframes. Frames with a difference of <0.5 are filtered out to ensure diversity (difference = 1 means completely different).
[0392] 3. Multi-objective optimization screening:
[0393] Construct a multi-objective optimization function:
[0394] Score = Semantic representativeness × 0.5 + Visual diversity × 0.3 + Emotional salience × 0.2
[0395] The optimal solution was obtained by using the Non-Dominated Sorting Genetic Algorithm (NSGA-II), which balances semantic, visual, and emotional dimensions, and finally selects the TOP-N keyframes (N = 5%-10% of the total number of frames).
[0396] In summary, the visual saliency calculation module outputs a saliency map that provides the visual foundation for sentiment fusion. The fused sentiment saliency map guides candidate frame extraction, and the screening module further optimizes the candidate set. If the semantic coverage of the keyframes after screening is insufficient (e.g., important events are missed), the system will backtrack to the candidate generation module and adjust the sampling strategy (e.g., reducing the uniform sampling interval), forming a closed loop of "saliency calculation → sentiment fusion → candidate generation → screening optimization". This process achieves a keyframe extraction accuracy of 92.4%, a 27% improvement over traditional saliency methods. Particularly in e-commerce promotional videos, the recall rate of product keyframes increased from 68% to 95%.
[0397] Optionally, the lens segmentation and reconstruction subunit includes:
[0398] The shot boundary detection module is used to detect shot transition points in semantically labeled materials to generate a list of shot boundary time points;
[0399] The shot semantic annotation module is used to annotate the shots corresponding to the shot boundary time point list with semantic theme and sentiment tags in order to generate a semantic shot set;
[0400] The editing decision mapping module is used to perform semantic mapping and matching processing between the semantic shot set and editing decision instructions and the timeline to generate a mapped shot set;
[0401] The temporal reorganization optimization module is used to optimize the narrative logic and reorganize the temporal sequence of the mapped shot set to generate a reorganized shot sequence.
[0402] When the shot boundary detection module detects shot transition points in semantically labeled materials and generates a list of shot boundary time points, it includes the following steps:
[0403] 1. Visual feature difference detection:
[0404] The dual-threshold inter-frame difference method is adopted: the difference in brightness histogram and pixel gradient between adjacent frames are calculated. When both exceed the threshold (brightness difference > 20% and gradient difference > 30%), they are marked as potential boundary points.
[0405] Gaussian pyramid multi-scale analysis is introduced to fuse frame difference results from different resolution layers (such as 1 / 4, 1 / 2, and original size) to avoid misjudgment at a single scale (such as false boundaries caused by fast motion).
[0406] 2. Semantic mutation detection:
[0407] The semantic similarity between adjacent frames is calculated using a pre-trained CLIP model. When the cosine distance between semantic vectors is greater than 0.6, it is determined to be a semantic abrupt change boundary (such as a scene switching from indoor to outdoor).
[0408] By combining entity changes in semantically labeled materials, if a sudden change is detected in the main entity (e.g., "lipstick" becomes "foundation") or scene label (e.g., "close-up" becomes "panorama"), it is forcibly marked as the lens boundary.
[0409] 3. Boundary point post-processing:
[0410] Sliding window smoothing (window size 5 frames) is used to eliminate dense false boundaries and preserve real boundary points; artificial heuristic boundaries are inserted for long shots (>30 seconds) (such as forced splitting every 20 seconds) to improve the flexibility of subsequent editing.
[0411] Output a list of boundary time points in the format `[00:00:05.234,00:00:12.567,...]`, accurate to the millisecond level.
[0412] The semantic annotation module annotates the lenses corresponding to the lens boundaries with semantic themes and sentiment tags, and generates a semantic lens set by including the following steps:
[0413] 1. Multimodal feature extraction:
[0414] Visual features: Use Swing Transformer to extract global semantic features from keyframes of the shot (such as "product display" and "promotional scene"), and combine it with YOLOv8 to detect entity categories (such as "lipstick" and "discount tag").
[0415] Audio features: Extract MFCC features and analyze intonation changes using BiLSTM to identify emotional tendencies such as "excitement" and "rapidity";
[0416] Text features: Use BERT to encode semantic vectors for OCR text (such as promotional slogans) in the footage.
[0417] 2. Semantic topic classification:
[0418] A three-layer classifier was constructed: the first layer identifies basic scenes (indoor / outdoor / close-up), the second layer identifies event types (product display / promotional information), and the third layer identifies specific entities (such as "lipstick swatches" and "limited-time discounts"). The classifier was fine-tuned based on over 100,000 labeled lenses, achieving an F1 score of 0.92.
[0419] Knowledge graph enhancement is introduced: the identified entities (such as "lipstick") are associated with the "beauty products" node in the knowledge graph to enrich the semantic hierarchy (such as "lipstick → beauty → fast-moving consumer goods").
[0420] 3. Sentiment tag generation:
[0421] It integrates visual emotions (facial expressions, colors) and audio emotions (tone, rhythm), and uses an attention mechanism to weight the output (60% visual and 40% audio) to output 8 types of emotional labels and intensity values (0-1), such as "excitement" and "warmth".
[0422] Output semantic shot structure:
[0423]
[0424]
[0425] The editing decision mapping module maps and matches semantic shots with editing decision instructions to generate a set of mapped shots, including the following steps:
[0426] 1. Decision instruction analysis:
[0427] Extract timeline segments (opening / main / ending), key events (such as "product close-up" and "special offer"), and emotional requirements (such as "fast pace" and "high excitement") from the editing decision instructions to construct a target semantic vector (such as "main body + product display + excitement").
[0428] 2. Semantic matching algorithm:
[0429] A three-layer matching strategy is adopted:
[0430] Basic matching: Calculate the keyword overlap rate between the semantics of the shot and the decision instructions (such as the matching degree of "product display");
[0431] Emotional matching: Calculate the difference between the emotional intensity of the shot and the decision requirements (e.g., if the excitement level is required to be >0.6, select the corresponding shot);
[0432] Structure matching: Constraints based on timeline segments (e.g., only "brand logo" type shots are matched in the opening segment).
[0433] The Transformer encoder is used to encode shot semantics and decision instructions into a unified spatial vector, which is then sorted by cosine similarity (threshold > 0.6).
[0434] 3. Mapping optimization mechanism:
[0435] When dealing with conflict situations (such as high semantic matching but insufficient sentiment), reinforcement learning reordering is enabled: prioritize retaining sentiment-matching shots (weight 0.7), and then consider semantics (0.3).
[0436] Output mapping results:
[0437]
[0438] The temporal reassembly optimization module performs narrative logic optimization and temporal reassembly on the mapped shots, generating a reassembled shot sequence, including the following steps:
[0439] 1. Narrative Logic Modeling:
[0440] Construct an advertising narrative map and define narrative patterns such as "problem-solution" and "feature-advantage-benefit". Each pattern corresponds to the temporal constraints of the shot type (e.g., a "product problem" shot must be followed by a "solution" shot).
[0441] The narrative flow is represented by a state machine, where the state is the semantic category of the shot, and the transition condition is the semantic relevance (e.g., "product display" can transition to "usage effect" or "promotional information").
[0442] 2. Timing optimization algorithm:
[0443] A shot sequence graph is constructed based on a graph neural network (GNN), where nodes represent shots and edge weights represent semantic coherence (calculated using BERT), visual fluency (calculated using optical flow to measure motion differences), and emotional consistency (absolute value of the difference).
[0444] The optimal sequence is solved using the simulated annealing algorithm. The objective function is:
[0445] Maximize (narrative logic compliance × 0.4 + visual fluency × 0.3 + emotional consistency × 0.3)
[0446] When processing long videos, first optimize locally by narrative segment (e.g., every 10 seconds), and then make global fine-tuning to improve computational efficiency.
[0447] 3. Recombinant sequence generation:
[0448] Output a reconstructed sequence including shot order and transition methods:
[0449]
[0450] In summary, the shot boundary detection module defines the shot range by outputting boundary points, the semantic annotation module assigns semantic and emotional attributes to each shot, the mapping module aligns the shots with decision instructions, and the reorganization and optimization module generates the final sequence. If a narrative logic break is found during reorganization (such as an irrelevant shot following "promotional information"), the system will backtrack to the mapping module to re-select shots, forming a closed loop of "boundary detection → semantic annotation → mapping matching → reorganization and optimization". This process improves the narrative rationality of shot reorganization by 35% and increases user click-through rate by 28% compared to random arrangement. In particular, in e-commerce promotional videos, optimizing the display order of product conversion-related shots increases conversion rate by 19%.
[0451] Optionally, the automatic generation subunit for transition effects includes:
[0452] The transition type prediction module is used to perform semantic similarity and visual difference analysis on adjacent shots in the recombined shot sequence to generate transition type prediction results.
[0453] The special effects parameter generation module is used to generate special effects parameters (duration / direction / effect style) based on the transition type prediction results and editing decision instructions, so as to generate a set of special effects parameters;
[0454] The transition effects generation module is used to instantiate and render transition effects from a set of effect parameters to generate transition effect materials.
[0455] The transition smoothness optimization module is used to optimize the temporal smoothness and visual coherence of transition effects materials and recombined shot sequences to generate shot sequences with transitions.
[0456] The transition type prediction module performs semantic similarity and visual difference analysis on adjacent shots in the reconstructed shot sequence and generates transition type prediction results, including the following steps:
[0457] 1. Multimodal feature extraction:
[0458] Visual features: The CLIP model is used to extract visual semantic features of keyframes of adjacent shots, and visual similarity (such as cosine distance) is calculated by comparing feature vectors; at the same time, the optical flow method is used to calculate the difference in motion vectors between frames to quantify the degree of visual abruptness of shot switching.
[0459] Semantic features: Extract topic tags (such as "product display" and "promotional information") and sentiment intensity from shot semantic annotations, and calculate semantic coherence (such as whether the topics of adjacent shots are related) through BERT.
[0460] 2. Transition Type Classification Model:
[0461] A three-layer classifier is constructed: the first layer identifies the broad category of transition types (shear / dissolve / slide / effect), the second layer subdivides the specific types (e.g., "fast cut," "fade in / fade out," "left / right slide"), and the third layer optimizes parameter preferences (e.g., the direction preference of effect transitions). The model is trained on over 500,000 labeled camera pairs, achieving an F1 score of 0.91.
[0462] An attention mechanism is introduced to dynamically adjust feature weights based on shot type: visual difference weights are added between close-up shots of the product (60%), and semantic weights are added between shots of scene transitions (70%).
[0463] 3. Prediction results generation:
[0464] Output transition type and confidence level, such as: `{"type":"cut","confidence":0.85}` (fast cut), `{"type":"fade","confidence":0.72}` (fade in / fade out), supporting "automatic" mode (determined by the model) or "manual" mode (overridden according to user preference).
[0465] When the special effects parameter generation module generates a set of special effects parameters (duration / direction / style) based on the transition type prediction results and editing decision instructions, it includes the following steps:
[0466] 1. Parameter rule engine:
[0467] A transition parameter knowledge base has been built, containing 200+ preset rules:
[0468] Quick cut: Fixed duration 0.1s, no direction, no default style;
[0469] Fade in / out: Duration dynamically adjusts according to emotional intensity (high emotion → 0.5s, low emotion → 1s), direction defaults to "none";
[0470] Slide: The direction is the same as the direction of camera movement (detected by optical flow method), and the duration is positively correlated with the duration of the shot (shot length → 1.2s).
[0471] 2. User intent adaptation:
[0472] The parameters in the editing decision instructions (such as "fast pace" and "technological effects") are analyzed, and the parameters are adjusted through reinforcement learning.
[0473] Under the demand for "fast pace", the duration of all transitions has been shortened by 20%;
[0474] For a "technological feel," prioritize styles such as "matrix switching" and "particle effects," and set the direction to "diagonal."
[0475] 3. Parameter conflict resolution:
[0476] When handling conflicting parameters (such as "slow pace" versus "fast cut"), prioritize them as follows: user command > emotional intensity > visual features. For example, if the user insists on a "fast cut", ignore the pace matching and only adjust the duration to a compromise value of 0.3 seconds.
[0477] Output parameter set: `{"duration":0.5s,"direction":"left_to_right","style":"glow_transition"}`.
[0478] The transition effects generation module instantiates and renders the effect parameter set, and generates transition effect materials through the following steps:
[0479] 1. Calling special effects template library:
[0480] We maintain 100+ transition effect templates, categorized by type (such as basic transitions, e-commerce effects, and technology effects). Each template includes pre-rendered resources and configurable parameter interfaces. For example, the "Golden Glitter" effect template supports adjusting parameters such as light intensity and particle density.
[0481] 2. Real-time rendering engine:
[0482] GPU-accelerated real-time rendering frameworks (such as lightweight versions of Unity / Unreal Engine) dynamically generate effects after receiving parameters:
[0483] For fade-in / fade-out transitions, calculate the pixel blending weight curves of adjacent shots;
[0484] For the "slide" transition, generate a layer displacement animation with motion blur;
[0485] For special effects transitions, instantiate a 3D model (such as a rotating cube) and bind parameters (rotation speed, material color).
[0486] 3. Material Generation Strategy:
[0487] The system employs a "pre-rendering + real-time compositing" mode: common transitions are pre-rendered as alpha channel video segments (such as "fast cut" and "fade in / fade out"), while special effects are generated in real time. The generated transition footage has the same resolution as the original video, and the frame rate matches the project settings (such as 30fps).
[0488] When optimizing the transition smoothness of transition effects footage and reassembled shot sequences for temporal smoothness and visual coherence, the transition smoothness optimization module includes the following steps:
[0489] 1. Timing smoothing:
[0490] Optical flow frame interpolation: For dynamic transitions such as "sliding" and "scaling", the motion vectors of adjacent shots are calculated to generate intermediate transition frames (such as inserting 2 frames every 0.1s) to eliminate the sense of stuttering.
[0491] Audio-video synchronization: Analyze the changes in audio energy before and after the transition, and adjust the transition duration to synchronize the audio transition with the visual transition (such as completing the transition at the beat of the music drum).
[0492] 2. Visual consistency optimization:
[0493] Color matching: Use 3D LUTs to unify the color tones of shots before and after transitions to avoid abrupt color changes (such as adding a gradient filter when suddenly switching from cool to warm tones).
[0494] Brightness balance: Histogram matching ensures consistent brightness distribution before and after transitions, eliminating visual flicker (e.g., adding brightness transition effects when transitioning from a dark scene to a bright scene).
[0495] 3. Smoothness assessment and correction:
[0496] Transition quality is evaluated using the VMAF-FR (Full Reference) index. If the score is less than 85, parameters are automatically adjusted.
[0497] Lag → Increase frame interpolation;
[0498] Color contrast → Enhances LUT matching;
[0499] Audio and video out of sync → Recalculate the audio synchronization point.
[0500] In summary, the transition type prediction module outputs type-guided parameter generation, which in turn drives effects rendering, while the optimization module ensures smoothness. If visual inconsistencies persist after optimization (such as excessive color differences), the system will backtrack to the parameter generation module to adjust the LUT parameters, forming a closed loop of "prediction → parameters → generation → optimization." This process improves transition effect generation efficiency by 40% and increases user satisfaction by 25% compared to manual production. Particularly in e-commerce promotional videos, the smoothness of fast-paced transitions is significantly improved, increasing view completion rates by 18%.
[0501] Optionally, the multitrack audio fusion subunit includes:
[0502] The audio feature extraction module is used to extract multi-track features of speech, sound effects and background music from the shot sequence with transitions to generate an audio feature set.
[0503] The audio scene matching module is used to perform audio scene (promotion / warm / dynamic) matching processing based on the audio feature set and semantic parsing results to generate a set of candidate audio materials;
[0504] The audio parameter adjustment module is used to adjust the volume, rhythm, and semantic synchronization of the candidate audio material set to generate adjusted audio material;
[0505] The multitrack mixing and rendering module is used to perform multitrack audio mixing and audio-visual synchronization rendering on the adjusted audio material and the shot sequence with transitions to generate a preliminary montage video.
[0506] The audio feature extraction module extracts multi-track features of speech, sound effects, and background music from a sequence of shots with transitions to generate an audio feature set. This process includes the following steps:
[0507] 1. Multi-track separation technology:
[0508] The Wave-U-Net network is used to separate the original audio, decomposing the input mixed audio (such as human voice + background music + ambient sound) into independent speech tracks, background music tracks, and sound effect tracks. The network learns the spectral characteristics of different audio sources through training, achieving a separation accuracy of 90% (SDR≥15dB).
[0509] Time-frequency domain feature extraction is performed on each separated track:
[0510] Speech track: Extract pitch, speech rate, energy envelope, and Mel-frequency cepstral coefficients (MFCC);
[0511] Background music track: Extract rhythm (BPM), key, and emotional tags (such as "upbeat" or "soothing");
[0512] Audio track: Extract sound pressure level (SPL), spectral distribution, and duration.
[0513] 2. Audio Scene Analysis:
[0514] By combining shot semantic annotations (such as "product display" and "promotional information"), each audio segment is classified into scenarios. For example, when the shot semantics are detected as "limited-time discount", the focus is on analyzing keywords in the audio track (such as "buy now") and their audio energy changes.
[0515] An audio emotion classifier was built, which uses the ResNet architecture to identify emotions (such as "excitement" and "calm") in audio segments, achieving a classification accuracy of 85%.
[0516] 3. Feature set generation:
[0517] Output a set of structured features, for example:
[0518]
[0519]
[0520] The audio scene matching module performs audio scene matching based on the audio feature set and semantic parsing results, and generates a candidate audio material set by including the following steps:
[0521] 1. Scene Template Library:
[0522] Build a knowledge base containing 100+ audio scene templates, each template associated with a specific combination of audio features:
[0523] Promotional scenarios: High-energy background music (BPM≥120), clear voice (SNR≥15dB), high-frequency sound effects (such as countdown prompts);
[0524] Warm and inviting atmosphere: soothing background music (BPM≤80), gentle voice, and low-frequency ambient sounds (such as soft wind).
[0525] Dynamic scenes: strong rhythmic drum beats (BPM≥130), dynamic voice changes, and high-frequency impact sound effects.
[0526] 2. Similarity matching algorithm:
[0527] Calculate the similarity between the current audio features and the scene template using a weighted distance metric:
[0528] The speech feature weight is 0.4 (keyword matching degree + pitch stability);
[0529] Background music weighted at 0.4 (BPM matching degree + emotional consistency);
[0530] Sound effect weight 0.2 (type matching degree + time synchronization).
[0531] 3. Candidate material screening:
[0532] Retrieve the candidate audio with the highest matching degree from the audio library:
[0533] For promotional scenarios, prioritize music that includes "promotional horn" sound effects and upbeat rhythms;
[0534] For warm and inviting scenes, select background music that primarily features piano or strings;
[0535] For dynamic scenes, match electronic dance music or rock music.
[0536] Output candidate set:
[0537]
[0538] The audio parameter adjustment module adjusts the volume, rhythm, and semantic synchronization of candidate audio materials. When generating the adjusted audio material, it includes the following steps:
[0539] 1. Volume balance control:
[0540] Dynamically adjust the volume of each track based on Voice Activity Detection (VAD):
[0541] When there is valid speech in the audio track, the background music is automatically reduced by 3-5dB;
[0542] When keywords (such as "attention" or "look here") are detected, temporarily increase the volume of the audio track by 2dB;
[0543] The sound track dynamically adjusts its gain based on importance levels (e.g., "promotional horn" has higher priority than regular click sounds).
[0544] 2. Rhythm synchronization optimization:
[0545] Analyze video transition timing (such as camera cuts) and align the background music's drum beats or melodic climaxes with the transitions. For example, when a fast transition is detected, ensure the music has a rhythmic accent at the moment of transition.
[0546] Time-stretch the audio track and adjust the speech rate without changing the pitch to match the rhythm of the video (e.g., a speech rate of ≥160 words / minute corresponds to a fast-paced video).
[0547] 3. Enhanced semantic synchronization:
[0548] Adjust audio parameters according to the semantics of the shot:
[0549] In "product close-up" shots, enhance ambient sound effects (such as product packaging sounds);
[0550] In the "Special Offers" shot, add low-frequency emphasis sounds (such as bass drums);
[0551] In emotionally charged scenes, increase the dynamic range of the background music (e.g., from soft to loud crescendo).
[0552] The multi-track blending and rendering module performs multi-track blending and synchronized audio-visual rendering on the adjusted audio material and the shot sequence with transitions to generate a preliminary montage video, including the following steps:
[0553] 1. Multitrack mixing engine:
[0554] Using 3D audio spatialization technology, virtual spatial positions are assigned to different audio tracks:
[0555] The audio track is positioned directly in front (0° azimuth) at a distance of 0.5m;
[0556] Background music tracks are distributed in a 360° surround pattern, 2m apart;
[0557] Special sound effects are positioned according to semantics (e.g., the sound effect for "product on the left" is placed in the left channel).
[0558] 2. Audio-visual synchronization technology:
[0559] The motion speed between video frames is calculated using optical flow, and the audio playback rate is adjusted synchronously. For example, when the video is played at an accelerated speed, the audio is time-compressed using the WSOLA algorithm while maintaining the pitch.
[0560] For transition effects (such as fade-in and fade-out), apply audio fade-in and fade-out effects simultaneously to ensure a consistent visual and auditory transition.
[0561] 3. Real-time rendering and quality control:
[0562] Multitrack audio transmission is performed using the AES67 standard to ensure low-latency (<1ms) mixing;
[0563] The output audio is dynamically compressed (DRC), limiting peak values (-3dBFS) and increasing average volume (LUFS = -16).
[0564] Generate a preview version and evaluate speech clarity using the Short-Time Objective Intelligibility Index (STOI). If the STOI < 0.9, automatically adjust the speech enhancement parameters.
[0565] In summary, the audio feature extraction module provides the data foundation for scene matching, the matching results guide parameter adjustments, and the adjusted audio is then mixed, rendered, and synchronized with the video. If audio-visual asynchrony is found after rendering (such as misalignment between speech and lip movements), the system will backtrack to the parameter adjustment module to correct the time offset, forming a closed loop of "extraction → matching → adjustment → rendering". This process improves audio production efficiency by 50% and achieves an audio-visual synchronization accuracy of 98%. Especially in e-commerce videos, improved speech clarity increases viewer comprehension by 22% and product conversion rate by 15%.
[0566] Optionally, the effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video, specifically including the following steps:
[0567] The visual effects enhancement subunit is used to perform color correction, sharpening, and special effects overlay processing on the initial montage video to generate a visually enhanced video;
[0568] The intelligent subtitle generation and positioning subunit is used to perform speech recognition and subtitle generation on visually enhanced videos, and to perform subtitle positioning processing in combination with screen area detection to generate videos with subtitles.
[0569] The multi-terminal adaptation and optimization subunit is used to perform multi-terminal adaptation and optimization processing on subtitled videos in terms of resolution, bit rate and format, so as to generate adapted and optimized videos.
[0570] The quality assessment and output subunit is used to evaluate the smoothness, image quality, and semantic consistency of the adapted and optimized video, and to perform optimization processing based on the evaluation results to generate the target mashup video.
[0571] Optionally, the visual effects enhancement subunit includes:
[0572] The color space conversion module is used to perform color space (RGB / YCbCr) conversion and color gamut normalization on the initial mixed-edit video to generate a color-normalized video;
[0573] The color correction module is used to adjust the brightness / contrast, balance the color, and apply stylized filters to color-normalized videos to generate color-corrected videos.
[0574] The super-resolution module is used to enhance the resolution and restore the details of color-corrected videos to generate super-resolution videos.
[0575] The visual effects overlay module is used to overlay dynamic stickers, special effects elements, and subtitle backgrounds on super-resolution videos to generate visually enhanced videos.
[0576] The color space conversion module performs color space (RGB / YCbCr) conversion and color gamut standardization on the initial montage video, generating a color-standardized video, including the following steps:
[0577] 1. Color Space Conversion Engine:
[0578] Matrix transformations are used to achieve bidirectional conversion between RGB and YCbCr color spaces. For example, the following formula can be used to convert RGB to YCbCr:
[0579] Y = 0.299R + 0.587G + 0.114B
[0580] Cb = -0.1687R0.3313G + 0.5B + 128
[0581] Cr = 0.5R0.4187G0.0813B+128
[0582] It supports conversion parameters for different standards such as BT.601 and BT.709, and automatically selects the appropriate parameter based on the target video platform (e.g., BT.709 for TVs and sRGB for mobile devices).
[0583] 2. Color gamut standardization processing:
[0584] Detect the color gamut space of the input video (such as Adobe RGB, DCI-P3), and compress it to the target color gamut (such as sRGB) using a color gamut mapping algorithm:
[0585] For colors that exceed the target color gamut, use saturation compression (keep the hue unchanged and reduce the saturation);
[0586] For HDR videos, use PQ or HLG curves for SDR mapping to preserve luminance information while avoiding overexposure.
[0587] 3. Output control:
[0588] When generating color-standardized videos, the color semantic information of the original video is preserved. For example, in e-commerce videos, it is ensured that the product color (such as the red of lipstick) is displayed consistently on different devices, with a color difference ΔE ≤ 3.
[0589] The color correction module performs brightness / contrast adjustments, color balance, and stylized filter processing on the color-normalized video. The process of generating a color-corrected video includes the following steps:
[0590] 1. Adaptive brightness and contrast adjustment:
[0591] Local histogram equalization is used to divide the video frame into 8×8 pixel blocks, calculate the histogram for each block and stretch it to the target range to improve dark details (such as the texture in the shadows of the product);
[0592] The brightness curve is dynamically adjusted through Gamma correction. For example, in e-commerce videos, the Gamma value of the product area is increased to 1.8 to make the product brighter.
[0593] 2.3D color balance processing:
[0594] 3D LUTs (lookup tables) are constructed for color mapping, supporting the following adjustments:
[0595] Skin tone optimization: Detect the face area and map the YCrCb value of the skin tone range to a more natural range;
[0596] Product color enhancement: Increase the saturation of the red channel by 20% in the lipstick area of beauty videos;
[0597] 10+ industry-specific LUTs are pre-set (such as "promotional warm colors" and "tech cool colors"), which can be automatically applied through semantic parsing.
[0598] 3. Stylized filter generation:
[0599] Custom filters can be generated based on GAN networks. For example, if the keyword "cinematic" is input, the model will automatically generate filters with vignetting and film grain.
[0600] It supports real-time preview and parameter fine-tuning, such as adjusting filter intensity (0-100%) and application range (full frame / local).
[0601] The super-resolution module enhances the resolution and restores the details of color-corrected videos. The process of generating a super-resolution video includes the following steps:
[0602] 1. Multi-frame super-resolution model:
[0603] An improved Real-ESRGAN model is used, taking 5 consecutive video frames as input, and aligning the motion between frames using optical flow to improve resolution by utilizing temporal information.
[0604] Use single-frame super-resolution for static frames (e.g., upscaling 1080P to 4K);
[0605] Multi-frame fusion is used for dynamic frames to reduce motion blur (such as jagged edges when a person is moving).
[0606] 2. Detail recovery algorithm:
[0607] Introducing an edge-aware loss function preserves edge details while improving resolution:
[0608] Apply a sharpening filter to the product outline (such as the edge of a lipstick tube) to enhance edge contrast;
[0609] For textured areas (such as cloth), use wavelet transform to decompose high-frequency details and then re-blend them.
[0610] 3. Real-time processing optimization:
[0611] A progressive super-resolution strategy is adopted, first performing 2x super-resolution, and then performing 4x super-resolution on key areas (such as product close-ups) to balance image quality and performance, with a processing speed of 20fps (4K output).
[0612] The visual effects overlay module applies dynamic stickers, special effects elements, and subtitle backgrounds to super-resolution videos to generate visually enhanced videos, including the following steps:
[0613] 1. Smart sticker positioning system:
[0614] By combining YOLOv8 to detect target locations (such as lipstick or faces) in videos, dynamic stickers automatically attach to the target areas.
[0615] For "limited-time discount" stickers, prioritize placing them in blank areas of the detection screen (such as the upper right corner);
[0616] For product stickers, attach them close to the edge of the product (e.g., 20 pixels above the lipstick).
[0617] 2. Dynamic effects generation engine:
[0618] Animation effects are generated based on a timeline, supporting the following types:
[0619] Entry animation: The sticker flies in from off-screen, lasting 0.5 seconds;
[0620] Emphasis animation: Promotional label flashes (frequency 2Hz) for 3 seconds;
[0621] Alpha channel blending technology is used to ensure that the special effects blend naturally with the original video, with an edge feathering radius of 2 pixels.
[0622] 3. Optimized subtitle background texture:
[0623] Detect text areas on the screen (using OCR) to prevent subtitles from obscuring the original text;
[0624] Automatically generate semi-transparent backgrounds (such as black semi-transparent rectangles) to improve the readability of subtitles. The transparency of the background is automatically adjusted according to the background brightness (bright background → 30% transparency, dark background → 70% transparency).
[0625] In summary, the color space conversion module ensures color consistency, providing a standard input for subsequent corrections; the color correction module enhances visual expressiveness; the super-resolution module improves image quality details; and finally, the effects overlay module adds elements required for the business. If color deviations are found after overlaying effects (such as excessive color difference between stickers and backgrounds), the system will backtrack to the color correction module for readjustment, forming a closed loop of "conversion → correction → super-resolution → overlay". This process improves video visual quality by 40%, achieves 98% product color accuracy in e-commerce scenarios, and increases effects overlay efficiency by 10 times compared to manual processing.
[0626] Optionally, the intelligent subtitle generation and positioning subunit includes:
[0627] The speech recognition module is used to process the human voice track in the visually enhanced video by converting speech to text and adding punctuation to generate speech-text content;
[0628] The subtitle style generation module is used to generate subtitle style schemes by processing the font, color and animation effects of subtitles based on the audio text content and video semantics.
[0629] The image region analysis module is used to detect and process text regions, salient regions, and safe regions in visually enhanced videos to generate image region maps.
[0630] The subtitle positioning optimization module is used to plan the subtitle position and handle conflicts based on the audio text content, subtitle style scheme, and screen area map to generate videos with subtitles.
[0631] The speech recognition module performs speech-to-text conversion and adds punctuation to the human voice track in the visually enhanced video. When generating the speech-text content, the following steps are included:
[0632] 1. Multimodal speech processing framework:
[0633] The Whisper-large-v3 model is used as the basic speech recognition engine, supporting Chinese and English bilingual languages and 10+ dialects (such as Cantonese and Sichuanese). It learns the mapping relationship between audio features and text through pre-training, and the word error rate (WER) is ≤5%.
[0634] Pre-amplifier audio noise reduction: Wave-U-Net is used to remove environmental noise (such as background noise and current noise) to improve speech purity. The SNR after noise reduction is ≥20dB.
[0635] 2. Dynamic punctuation addition algorithm:
[0636] Based on language models (such as GPT-2), the semantic structure of speech text is analyzed, and punctuation marks such as commas and periods are automatically inserted into long sentences.
[0637] A comma is inserted when a pause in speech is detected (a sudden drop in audio energy > 3dB);
[0638] Add a period after recognizing a complete semantic unit (such as "Buy Now").
[0639] Supports custom punctuation styles (e.g., using more exclamation marks in e-commerce promotions, and using strict punctuation in science videos).
[0640] 3. Real-time error correction mechanism:
[0641] Build an industry terminology database (such as "lipstick" and "promotion") and perform post-processing on the recognition results:
[0642] If "lipstick shade" is misidentified as "good lipstick", it will be automatically corrected based on semantic similarity (≥0.8);
[0643] Output text content with timestamps, for example:
[0644]
[0645] The subtitle style generation module generates subtitle style schemes based on the audio text content and video semantics, including fonts, colors, and animation effects, and includes the following steps:
[0646] 1. Semantic-style mapping engine:
[0647] Build a style knowledge base and define style rules for 20+ scenarios:
[0648] Promotional scenario: Bold white text with yellow outline (#FFFFFF+#FFD700), font selection "FZ Bold Simplified", animation "zoom in";
[0649] Warm and inviting scene: handwritten white text (#FFFFFF), 70% opacity, with a fade-in / fade-out animation;
[0650] BERT is used to parse the text sentiment (e.g., "limited-time offer" → excitement) and match it with the corresponding style template.
[0651] 2. GAN style generation technology:
[0652] Train a generative adversarial network (GAN) to generate custom styles. Input keywords (such as "technological feel") to generate blue subtitles with a glowing effect (#00BFFF+outer glow).
[0653] Supports real-time adjustment of parameters: font size (36-72px), stroke width (2-5px), and animation duration (0.3-1s).
[0654] 3. Dynamic style adaptation:
[0655] The style is automatically adjusted according to the text length: long texts are reduced in font size (e.g., from 48px to 36px if it exceeds 3 lines) to maintain readability;
[0656] Output style scheme:
[0657]
[0658] The video region analysis module performs text region, salient region, and safe region detection on the video, and generates a video region map, including the following steps:
[0659] 1. Multi-region detection network:
[0660] Text region: Using the DBnet+CRNN combined model, printed and handwritten text in the image is detected with pixel-level positioning accuracy, and it supports mixed Chinese, English and numbers text;
[0661] Saliency region: Using BASNet+ attention mechanism, visual focus (such as human face, product subject) is identified and saliency probability map (0-1 value) is output;
[0662] Safe Zone: Define 20% of the screen edge as the safe zone (to avoid subtitles being obscured), and 60% of the center as the critical zone (to prioritize content preservation).
[0663] 2. Regional conflict detection:
[0664] Construct a region conflict matrix and calculate the overlap rate between text regions and salient regions:
[0665] If the candidate subtitle position overlaps with the product close-up area by more than 30%, it is marked as a conflict.
[0666] Positions outside the safe zone are automatically excluded (e.g., no subtitles are placed in the top 10% area).
[0667] 3. Regional map generation:
[0668] Output a structured graph, for example:
[0669]
[0670]
[0671] The subtitle positioning and optimization module integrates voice text, style schemes, and screen area maps to plan subtitle positions and avoid conflicts, including the following steps:
[0672] 1. Greedy positioning algorithm:
[0673] Prioritize locations within safe zones, filtering in the following order:
[0674] 1. Safe area at the bottom of the screen (height 100-200px);
[0675] 2. Safe area at the top of the screen;
[0676] 3. Safety zones on both sides (100-200px wide);
[0677] Automatically wrap long text (≤15 characters per line) and set the line spacing to 1.2 times the font size.
[0678] 2. Conflict Avoidance Strategies:
[0679] If the candidate position overlaps with the text area, move it up / down 50px and try again;
[0680] If the overlap with a prominent area is greater than 20%, adjust the subtitle transparency (reduce to 50%) or add a semi-transparent background (black, 70% transparency) to improve readability.
[0681] 3. Real-time preview optimization:
[0682] Generate a preview frame with subtitles, evaluate the integration of subtitles with the screen using the SSIM metric, and readjust the position if SSIM < 0.8;
[0683] Output video with subtitles, and synchronize subtitle position information to metadata, for example:
[0684]
[0685] In summary, the text output by the speech recognition module drives style generation and localization planning, while image area analysis provides the basis for obstacle avoidance. If, after localization, it is found that subtitles obscure key content (such as product labels), the system will backtrack to the style generation module to adjust the font size or transparency, forming a closed loop of "recognition → style → analysis → localization". This process improves subtitle generation efficiency by 80%, reduces the obscuring conflict rate from 35% to 5%, and in e-commerce videos, improves subtitle readability by 27% and increases user click-through conversion rate by 19%.
[0686] Optionally, the multi-terminal adaptation optimization subunit includes:
[0687] The terminal parameter parsing module is used to parse and process the resolution, bitrate, and format requirements of the target output terminal (such as TikTok, YouTube, and TV) to generate a set of terminal parameters.
[0688] The resolution adaptation module is used to scale the resolution and remove black borders of videos with subtitles based on the set of terminal parameters in order to generate resolution-adapted videos.
[0689] The dynamic bitrate optimization module is used to perform bitrate and image quality balance optimization on resolution-adapted videos to generate bitrate-optimized videos.
[0690] The format encapsulation module is used to encapsulate the bitrate-optimized video into terminal formats and adapt the encoding protocol to generate adapted and optimized videos.
[0691] The terminal parameter parsing module parses the resolution, bitrate, and format requirements of the target output terminal and generates a terminal parameter set, including the following steps:
[0692] 1. Terminal Feature Library Construction:
[0693] Maintain a parameter configuration library containing 100+ mainstream terminals, such as:
[0694] TikTok: Recommended resolution 720p / 1080p, bitrate 3-8Mbps, format MP4 (H.264+AAC);
[0695] YouTube: Supports 4K / 8K, with dynamic bitrate adjustment (4K requires 25-40Mbps), MP4 / WEBM format;
[0696] Smart TV: 1080p / 2160p resolution, bitrate 8-15Mbps, supports Dolby Vision / HDR10.
[0697] The parameter template is automatically matched based on the terminal ID. If "Douyin" is detected, the short video optimization configuration is loaded.
[0698] 2. Dynamic parameter acquisition:
[0699] For unknown terminals (such as custom players), obtain User-Agent information via HTTP requests and combine it with machine learning models to predict optimal parameters;
[0700] It supports user-defined parameter overriding, such as specifying the maximum bitrate for a specific platform.
[0701] 3. Parameter conflict handling:
[0702] When multiple parameters conflict (such as high resolution versus low bitrate), prioritize the key metrics:
[0703] Short video platforms prioritize resolution (720p and above), while live streaming platforms prioritize frame rate (≥30fps).
[0704] Output structured parameter set:
[0705]
[0706]
[0707] When the resolution adaptation module performs resolution scaling and black border processing on videos with subtitles based on terminal parameters, it includes the following steps:
[0708] 1. Intelligent scaling algorithm:
[0709] A bicubic interpolation algorithm is used for scaling, and Lanczos filtering is applied to reduce blurring in edge detail areas;
[0710] Supports non-uniform scaling when the ratio of the source video to the target resolution is greater than 10%.
[0711] Prioritize preserving the main elements of the video (through salience detection), such as people's faces and product areas;
[0712] Instead of directly stretching, intelligently crop the edge areas (e.g., crop 5% on each side).
[0713] 2. Black border handling strategy:
[0714] When the aspect ratio of the source video is inconsistent with that of the target video, black borders are dynamically generated:
[0715] Calculate the black border ratio (e.g., when adapting a 16:9 video to a 4:3 terminal, add 12.5% black borders at the top and bottom);
[0716] Add a gradient effect (from pure black to -10% brightness) to the black border area to enhance the visual appeal;
[0717] Supports recalculation of subtitle position to prevent subtitles from being obscured by black borders.
[0718] 3. Progressive resolution adjustment:
[0719] For high-resolution videos (such as 4K), multi-level downsampling (4K→2K→1080p) is used to reduce information loss;
[0720] The output resolution is adapted to the video, and the metadata includes scaling factors and black border parameters.
[0721] When the dynamic bitrate optimization module performs bitrate-quality balance optimization on resolution-adapted videos, it includes the following steps:
[0722] 1. Content-aware bitrate allocation:
[0723] The video frames are divided into 8×8 macroblocks, and the content complexity is analyzed using a CNN:
[0724] Allocate more bitrate (+20%) to high-motion areas (such as areas where people are moving quickly);
[0725] Reduce bitrate in static background areas (-15%);
[0726] An additional 10% bitrate is added to text areas (such as subtitles and product labels) to ensure clarity.
[0727] 2. Dynamic bitrate control algorithm:
[0728] VBR (Variable Bit Rate) encoding is used, combined with Rate-Distortion optimization:
[0729] Increase the peak bitrate (twice the average bitrate) for scene transition frames (detected through inter-frame differences);
[0730] Reduce the bitrate for smooth transition frames, while maintaining the average bitrate within ±10% of the target value;
[0731] Supports ABR (Adaptive Bitrate) stream generation and outputs multiple bitrate versions (e.g., 3Mbps / 5Mbps / 8Mbps).
[0732] 3. Quality assessment feedback:
[0733] Encoding quality is evaluated in real time using PSNR and SSIM metrics, and parameters are automatically adjusted when SSIM < 0.95.
[0734] Reduce the I-frame interval (from the default 12 frames to 6 frames) to improve random access performance;
[0735] A rate compensation mechanism is enabled to perform secondary encoding on low-quality areas.
[0736] When the format encapsulation module performs terminal format encapsulation and encoding protocol adaptation for bitrate-optimized videos, it includes the following steps:
[0737] 1. Multi-format encapsulation engine:
[0738] Supports 10+ container formats including MP4, WEBM, and TS, automatically selecting the appropriate format based on the terminal requirements.
[0739] MP4 is prioritized for mobile devices (broadly compatible), while WEBM is added as an alternative for web devices;
[0740] Optimize packaging parameters:
[0741] Adjust the position of the moov atom to the beginning of the file to shorten the video loading time (<1s);
[0742] Set an appropriate segment duration (4-6 seconds) to support HTTP Live Streaming (HLS).
[0743] 2. Encoding protocol adaptation:
[0744] Video encoding supports H.264, H.265, VP9, etc., and audio encoding supports AAC, OPUS, AC-3.
[0745] For iOS devices, H.264+AAC combination is forced; for Android devices, H.265 is prioritized to reduce bandwidth consumption.
[0746] For terminals that require high-quality audio (such as smart TVs), enable AC-3 5.1 channel encoding.
[0747] 3. Metadata Injection and Validation:
[0748] Automatically add key metadata:
[0749] Technical parameters such as video duration, resolution, and frame rate;
[0750] Subtitle track information (language, position), chapter markers (such as segmentation of e-commerce products);
[0751] Perform format validation before output to ensure the file conforms to terminal specifications (such as YouTube's file size limit).
[0752] In summary, terminal parameter parsing provides target configurations for subsequent modules, resolution adaptation ensures visual integrity, bitrate optimization balances bandwidth and image quality, and format encapsulation guarantees compatibility. If format incompatibility is found after encapsulation (e.g., a certain browser cannot play the video), the system will backtrack to the format encapsulation module to change the encoding protocol, forming a closed loop of "parsing → adaptation → optimization → encapsulation". This process increases the video playback success rate on various terminals to 99%, speeds up loading by 40%, and reduces bandwidth consumption by 35%. Especially in mobile weak network environments, the stuttering rate drops from 22% to 5%.
[0753] Optionally, the quality assessment and output sub-unit includes:
[0754] The objective image quality assessment module is used to calculate and process objective image quality indicators such as VMAF and PSNR for adapted and optimized videos in order to generate image quality assessment results.
[0755] The smoothness detection module is used to perform frame rate consistency and motion smoothness detection on the adapted and optimized video in order to generate smoothness evaluation results.
[0756] The semantic consistency verification module is used to perform consistency verification processing on the semantics of subtitles and audio-visual semantics of the adapted and optimized video in order to generate semantic evaluation results.
[0757] The optimization iteration and output module is used to perform iterative optimization processing based on the evaluation results of image quality, smoothness, and semantics, and output the final video.
[0758] The objective image quality assessment module calculates objective image quality metrics such as VMAF and PSNR for the adapted and optimized video, and generates image quality assessment results, including the following steps:
[0759] 1. Multi-indicator integrated evaluation framework:
[0760] Parallel computation of metrics such as VMAF (Multi-Method Evaluation Fusion for Video), PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity), and MS-SSIM (Multi-Scale Structural Similarity):
[0761] VMAF uses a Netflix open-source model to analyze the perceived quality of the human eye through CNN, with a score range of 0-100 (≥90 is considered excellent).
[0762] PSNR measures pixel-level error, with a threshold of 30dB (≥30dB indicates no significant distortion);
[0763] SSIM / MS-SSIM assesses image structural similarity with a threshold ≥0.95.
[0764] Weighted fusion results: VMAF weight 0.6, SSIM weight 0.3, PSNR weight 0.1, forming a comprehensive score.
[0765] 2. Region Adaptive Analysis:
[0766] The video frames are divided into a 16×16 grid, and the weight of high semantic regions (faces, text) is increased by an additional 20%.
[0767] Reduce the weight of transition areas (such as fade-in and fade-out) by 15% to avoid the transition effects affecting the overall score.
[0768] 3. Defect Detection Engine:
[0769] Identify coding defects such as blockiness, ambiguity, and ringing.
[0770] Block effects are detected by edge gradient abrupt changes; if block effects are present in more than 5% of the area, it is considered a defect.
[0771] Blur detection uses the Laplacian operator to calculate image sharpness; if the image is below a threshold, it is marked as blurry.
[0772] Output a detailed evaluation report:
[0773]
[0774] The smoothness detection module detects video frame rate consistency and motion smoothness, and generates smoothness evaluation results, including the following steps:
[0775] 1. Frame rate fluctuation analysis:
[0776] The actual frame rate is parsed from the video metadata, and the frame rate fluctuation factor (FVF) is calculated:
[0777] Collect 100 frames of samples and calculate the standard deviation of the time interval between each frame, with a threshold of ≤0.01 seconds;
[0778] Mark frame rate mutations (such as from 30fps to 15fps), with the number of mutations ≤ 2 times / minute.
[0779] 2. Motion smoothness assessment:
[0780] Motion coherence is evaluated by calculating motion vectors between adjacent frames using optical flow methods.
[0781] Calculate the distribution entropy of the motion vector length; the lower the entropy value, the smoother the motion.
[0782] For fast-moving scenes (such as sports videos), additional motion blur detection is performed to ensure that the blur level is ≤0.8px.
[0783] 3. Lag detection algorithm:
[0784] Detecting video playback lag (stuttering):
[0785] Frame freezes are detected based on changes in DCT coefficients, and if the freeze lasts for more than 0.5 seconds, it is considered a stutter.
[0786] Calculate the frequency of stuttering (times / minute) and the cumulative duration of stuttering, with thresholds of ≤1 time / minute and ≤0.5 seconds / minute, respectively;
[0787] Output smoothness report:
[0788]
[0789]
[0790] The semantic consistency verification module verifies the consistency of subtitle-image and audio-visual semantics. When generating semantic evaluation results, it includes the following steps:
[0791] 1. Subtitle-video synchronization detection:
[0792] Text is extracted from the video footage using OCR technology and compared with the subtitle text.
[0793] Calculate the edit distance (Levenshtein distance); a match rate of ≥90% is considered acceptable.
[0794] Precise matching of key information such as product name and price, allowing for zero error;
[0795] Detect the synchronization between the subtitle display time and the screen content:
[0796] When a character speaks on screen, the subtitles should appear within 0.3 seconds after the lip movement begins and disappear within 0.2 seconds after it ends.
[0797] 2. Audio-visual semantic alignment:
[0798] Analyze the correlation between audio keywords and video scenes:
[0799] When the audio mentions "the product on the left", the screen should switch to the product on the left within 3 seconds;
[0800] For emotional expressions (such as an excited tone), verify whether the video footage matches (such as a quick edit of a promotional scene);
[0801] Construct a cross-modal semantic similarity model and calculate the semantic correlation between audio and video frames using the CLIP architecture, with a threshold of ≥0.7.
[0802] 3. Logical coherence verification:
[0803] Detect the logical coherence of video content:
[0804] Analyze whether the camera transitions conform to narrative logic (e.g., from product demonstration to usage demonstration);
[0805] Verify the compatibility between transition effects and scene changes (e.g., gradient transitions are used for soothing scenes, while fast cuts are used for dynamic scenes);
[0806] Output semantic evaluation report:
[0807]
[0808] When the optimization and output module iterates and optimizes based on the evaluation results and outputs the final video, it includes the following steps:
[0809] 1. Multi-dimensional optimization decision engine:
[0810] Construct a decision tree model and formulate optimization strategies based on the evaluation results:
[0811] If VMAF < 85, increase the bitrate by 15% and re-encode;
[0812] If the stuttering frequency is greater than 1 time per minute, adjust the GOP structure (reduce the I-frame interval);
[0813] If the subtitle matching accuracy is less than 0.9, repeat the speech recognition and timeline calibration.
[0814] It supports multiple rounds of iteration, with re-evaluation after each round of optimization until all thresholds are met.
[0815] 2. Quality-efficiency balance algorithm:
[0816] Using the Pareto optimality principle, a balance is sought between quality improvement and computational resource consumption:
[0817] Optimization operations with a marginal benefit of less than 5% in improving image quality will automatically terminate;
[0818] Prioritize optimizing user-sensitive areas (such as faces and text), and appropriately reduce the quality of secondary areas;
[0819] It supports user-defined quality preferences (such as "ultra-high image quality" and "fast export").
[0820] 3. Final output processing:
[0821] Integrate all optimization results to generate the final video file:
[0822] Add metadata tags to quality metrics (such as VMAF, PSNR);
[0823] Embed the quality assessment report as an adjunct to the video;
[0824] Supports multiple output versions (such as different bitrates and resolutions) to meet the needs of different distribution channels;
[0825] When outputting the final video, thumbnails and preview clips are automatically generated to facilitate content review.
[0826] In summary, image quality assessment provides a technical quality benchmark, smoothness detection ensures a good viewing experience, and semantic verification ensures accurate information delivery. If an assessment identifies a problem (such as desynchronized subtitles), the system will backtrack to the relevant module (such as the subtitle positioning optimization subunit) for correction, forming a closed loop of "assessment → decision → optimization → reassessment". This process has increased the video quality compliance rate from 78% to 96%, reduced manual review workload by 60%, and decreased the return rate due to video quality issues by 23% in e-commerce scenarios.
[0827] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 2 As shown, an electronic device includes a memory and a processor, wherein the memory stores a computer-executable program, and the processor is configured to run the computer-executable program to perform the method in any of the above examples.
[0828] The above Figure 2 In the embodiments, an exemplary explanation of the technical processing procedures for each step can be found above. Figure 1 The records.
[0829] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims. The application, device, module, or unit described in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with a certain function.
[0830] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing this invention, the functions of each unit can be implemented in one or more software and / or hardware components.
[0831] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, this application, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0832] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (this application), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0833] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0834] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0835] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, a network interface, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0836] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0837] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0838] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, this application, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0839] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific transactions or implement specific abstract data types. This invention can also be practiced in distributed computing environments where transactions are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0840] The various embodiments in this specification are described in a progressive manner, with identical or similar parts between the embodiments referred to or substituted for each other. For the embodiments of this application, since they are basically similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0841] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims. The present application, device, module, or unit described in the above embodiments is specifically implemented by a computer chip or entity, or by a product having a certain function.
[0842] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, this application, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A video mixing and editing system, characterized in that, include: The material preprocessing unit is used to preprocess the original video material to generate standardized editing material; The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions; The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video; The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video in order to generate the target mashup video; The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions, specifically including: The video semantic understanding subunit is used to identify key content and perform semantic parsing on semantically labeled materials to generate semantic parsing results; The user intent parsing subunit is used to perform natural language understanding and intent extraction processing on the user's input editing requirements in order to generate editing intent instructions; The template matching and strategy generation subunit is used to perform editing template matching and strategy generation processing based on semantic parsing results and editing intent instructions in order to generate editing strategy schemes; The editing logic arrangement subunit is used to perform timeline logic arrangement and keyframe marking processing on the editing strategy scheme in order to generate editing decision instructions.
2. The video mixing and editing system according to claim 1, characterized in that, The material preprocessing unit is used to preprocess the original video material to generate standardized editing material, specifically including: The multi-source format standardization subunit is used to perform format-unified conversion on the original video footage to generate format-standardized footage; The spatiotemporal resolution unification subunit is used to normalize the temporal frame rate and spatial resolution of format-standardized materials in order to generate spatiotemporally standardized materials. The content noise filtering subunit is used to perform noise removal and abnormal frame filtering on spatiotemporally normalized material to generate denoised material. The metadata semantic annotation subunit is used to parse and semantically annotate metadata such as shooting parameters and scene information of the denoised material to generate semantically annotated material.
3. The video mixing and editing system according to claim 1, characterized in that, The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video, specifically including: The keyframe intelligent extraction subunit is used to perform visual saliency and sentiment peak detection processing on semantically labeled materials to generate a set of keyframes; The shot segmentation and reconstruction subunit is used to perform shot segmentation and temporal reconstruction processing on the keyframe set and editing decision instructions to generate a reconstructed shot sequence; The automatic transition effects generation subunit is used to intelligently match and generate transition effects for the recombined shot sequence in order to generate a shot sequence with transitions. The multitrack audio fusion subunit is used to perform multitrack fusion processing of background sound effects, voice narration and background music on the shot sequence with transitions to generate a preliminary montage video.
4. The video mixing and editing system according to claim 1, characterized in that, The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video, specifically including: The visual effects enhancement subunit is used to perform color correction, sharpening, and special effects overlay processing on the initial montage video to generate a visually enhanced video; The intelligent subtitle generation and positioning subunit is used to perform speech recognition and subtitle generation on visually enhanced videos, and to perform subtitle positioning processing in combination with screen area detection to generate videos with subtitles. The multi-terminal adaptation and optimization subunit is used to perform multi-terminal adaptation and optimization processing on subtitled videos in terms of resolution, bit rate and format, so as to generate adapted and optimized videos. The quality assessment and output subunit is used to evaluate the smoothness, image quality, and semantic consistency of the adapted and optimized video, and to perform optimization processing based on the evaluation results to generate the target mashup video.
Citation Information
Patent Citations
Short video editing method and system based on artificial intelligence
CN119031197A
Short video clip synthesis method based on fuzzy logic
CN119854573A