Video mixing and cutting system

By standardizing video footage, performing semantic analysis, and automating editing, the problems of insufficient format compatibility, efficiency, and adaptability in video mashup technology have been solved, achieving high-quality and diverse video editing effects.

CN120897100AActive Publication Date: 2025-11-04GUANGZHOU TAIDONG TECH CO LTD

Patent Information

Application Number
CN202510927212.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-04
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing video montage technology lacks the ability to understand the semantics of multi-source video materials, resulting in insufficient narrative coherence and visual expressiveness in the editing results. It also suffers from poor format compatibility, low editing efficiency, and insufficient terminal adaptability.

Method used

The system employs a material preprocessing unit for standardization, an intelligent editing decision unit for semantic analysis, an automated editing execution unit for efficient editing, and an effect optimization output unit for visual enhancement and format optimization to generate a high-quality target montage video.

Benefits of technology

It improves the format compatibility, editing efficiency, and terminal adaptability of multi-source video materials, and significantly enhances the narrative coherence and visual expressiveness of the editing results, meeting the needs of large-scale and personalized content production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897100A_ABST
    Figure CN120897100A_ABST
Patent Text Reader

Abstract

The invention provides a video mixing and cutting system, which comprises a material preprocessing unit used for preprocessing an original video material to generate a standardized edited material; the intelligent editing decision-making unit is used for carrying out semantic analysis and editing logic generation on the standardized editing materials so as to generate an editing decision-making instruction; the automatic editing execution unit is used for performing automatic editing processing on the standardized editing material according to the editing decision instruction so as to generate a preliminary mixed cutting video; and the effect optimization output unit is used for performing visual effect enhancement and format optimization on the preliminary mixed cutting video to generate a target mixed cutting video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of data processing, and in particular to a video mixing and editing system. BACKGROUND

[0002] With the explosive growth of short video content creation demand, e-commerce platforms, self-media operations and brand promotion scenarios have higher requirements for the automation and intelligentization of video mixing and editing. As a key technology for processing multi-source video materials into complete works, the efficiency and quality of video mixing and editing directly affect the cost and dissemination effect of content production. Currently, the video mixing and editing schemes on the market mainly rely on manual editing or semi-automatic tools based on fixed templates, which cannot meet the needs of rapid processing of massive materials and personalized creation.

[0003] The existing video mixing and editing technology usually adopts the technical idea of "template matching + simple parameter adjustment": the materials are spliced through pre-set editing templates, and the editing is completed by combining the transition effects and time length allocation specified by artificial. This kind of scheme often lacks semantic understanding ability of video content, and cannot intelligently process according to the visual features, emotional attributes and narrative logic of the materials. For example, when facing video materials of different scenes (such as e-commerce product display and brand story videos), the existing technology cannot automatically identify the key content in the materials and match appropriate editing strategies, resulting in obvious deficiencies in the narrative coherence and visual expressiveness of the edited results.

[0004] Further, the existing scheme has the following technical defects: first, the material preprocessing process lacks a standardized mechanism, and has poor compatibility for multi-source formats, resolutions and quality of materials, which easily leads to quality loss or format incompatibility problems in the edited video; second, the editing decision process relies on artificial experience or fixed rules, and cannot generate dynamic editing logic based on semantic information of the materials (such as scene type, object recognition result), making it difficult to adapt to diversified creation needs; third, in the automatic editing execution process, the links of shot segmentation, transition generation and audio fusion lack intelligent optimization, resulting in low editing efficiency and single effect; fourth, the output link lacks multi-terminal adaptation and quality optimization mechanism, and cannot be optimized according to the characteristics of the playing platform, affecting the final playing effect. These technical bottlenecks make it difficult for the existing video mixing and editing scheme to balance efficiency and quality when facing large-scale and personalized content production needs. SUMMARY

[0005] Therefore, embodiments of the present application provide a video mixing and editing system to at least partially solve the above problems.

[0006] A video mixing and editing system comprises:

[0007] a material preprocessing unit configured to preprocess original video materials to generate standardized editing materials;

[0008] An intelligent editing decision unit is configured to perform semantic analysis and editing logic generation on the standardized editing materials to generate editing decision instructions.

[0009] An automated editing execution unit is configured to perform automated editing processing on the standardized editing materials according to the editing decision instructions to generate a preliminary mixed and edited video.

[0010] An effect optimization output unit is configured to perform visual effect enhancement and format optimization on the preliminary mixed and edited video to generate a target mixed and edited video. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present application, and other drawings can be obtained by those skilled in the art according to these drawings.

[0012] Figure 1 A structural schematic diagram of a video mixed and edited system according to an embodiment of the present application.

[0013] Figure 2 A structural schematic diagram of an electronic device according to an embodiment of the present application is provided. DETAILED DESCRIPTION

[0014] In order to make the person skilled in the art better understand the technical solutions in the embodiments of the present application, the following will combine the drawings in the embodiments of the present application to clearly and detailedly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the embodiments of the present application, not all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those skilled in the art should belong to the scope of protection of the embodiments of the present application.

[0015] It should be understood that the terms "first", "second" and "third" and the like in the claims, specification and drawings of the present disclosure are used to distinguish different objects, not to describe a particular order. The terms "include" and "contain" used in the specification and claims of the present disclosure indicate the presence of the described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.

[0016] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the claims, the singular forms "a," "an" and "the" include plural referents unless the context clearly dictates otherwise. It is further to be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, and that the term "at least one of" is equivalent to the term "one or more of.

[0017] Figure 1 A structural schematic diagram of a video mixing and editing system according to an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, a video mixing and editing system includes: Figure 1

[0018] a material preprocessing unit configured to preprocess original video materials to generate standardized editing materials;

[0019] an intelligent editing decision unit configured to perform semantic analysis on the standardized editing materials and generate editing logic to generate editing decision instructions;

[0020] an automated editing execution unit configured to perform automated editing processing on the standardized editing materials according to the editing decision instructions to generate a preliminary mixed and edited video;

[0021] an effect optimization output unit configured to perform visual effect enhancement and format optimization on the preliminary mixed and edited video to generate a target mixed and edited video.

[0022] The video mixing and editing system according to the present application has the following technical advantages:

[0023] 1. The material preprocessing unit is configured to preprocess original video materials to generate standardized editing materials, which can effectively solve the compatibility problem of multi-source formats, resolutions and qualities, avoid loss of picture quality due to format incompatibility, provide a unified standard material basis for subsequent editing, and improve the overall quality and stability of the video.

[0024] 2. The intelligent editing decision unit is configured to generate editing logic based on semantic analysis of the standardized materials, breaking through the limitations of traditional reliance on artificial experience or fixed rules. The system can automatically identify semantic information such as material scene type and object features, dynamically generate adaptive editing strategies, significantly enhance the narrative coherence and visual expressiveness of the editing results, and meet diverse creative needs.

[0025] 3. The automated editing execution unit is configured to perform automated processing according to the intelligent generated editing decision instructions, and to realize intelligent operation in the aspects of shot segmentation, transition generation and audio fusion. Compared with the editing methods in the prior art which lack intelligent optimization, the editing efficiency is greatly improved, the problem of single effect is avoided, and high-quality and diversified editing effects are achieved.​

[0026] 4. The effect optimization output unit performs visual enhancement and format optimization on the preliminary mixed and cut video. In combination with multi-terminal adaptation technology, targeted adjustments can be made according to the characteristics of different playing platforms, ensuring that the video presents the best playing effect on various terminals, solving the problem of insufficient adaptability of existing solutions in the output link, effectively balancing efficiency and quality, and meeting the large-scale and personalized content production needs.

[0027] Optionally, the material preprocessing unit is configured to preprocess the original video material to generate standardized cut material, specifically including the following steps:

[0028] The multi-source format standardization subunit is configured to perform format uniform conversion processing on the original video material to generate format standardized material.

[0029] The space-time resolution uniformization subunit is configured to perform time frame rate and space resolution normalization processing on the format standardized material to generate space-time standardized material.

[0030] The content noise filtering subunit is configured to perform noise elimination and abnormal frame filtering processing on the space-time standardized material to generate denoising material.

[0031] The metadata semantic annotation subunit is configured to perform shooting parameter, scene information and other metadata analysis and semantic annotation processing on the denoising material to generate semantic annotation material.

[0032] Optionally, the multi-source format standardization subunit includes:

[0033] The format detection module is configured to perform format type and encoding protocol detection processing on the original video material to generate format description information.

[0034] The format conversion execution module is configured to perform uniform conversion processing of container format and encoding protocol according to the format description information to generate format conversion intermediate material.

[0035] The encoding optimization module is configured to perform rate control and encoding parameter optimization processing on the format conversion intermediate material to generate encoding optimized material.

[0036] The format verification module is configured to perform format integrity and protocol compatibility verification processing on the encoding optimized material to generate format standardized material.

[0037] Specifically, the format detection module adopts a "metadata first + deep learning assisted" dual-track detection strategy. First, read the header metadata of the video file, such as the ftyp (file type) and moov (movie structure) atoms in the MP4 format, and the RIFF (Resource Interchange File Format) block structure in the AVI format, to extract the container type (such as MP4, AVI, MOV) and basic encoding information (such as H.264, H.265). If the metadata is missing or damaged, enable the deep learning model: convert the first 5 seconds of the video into a fixed-size image sequence, input the pre-trained convolutional neural network (such as ResNet variant), which can identify the encoding feature pattern (such as H.264 macroblock structure, VP9 intra prediction mode) in the video frame by learning a large amount of labeled data (containing 100+ format samples), and then determine the encoding protocol type. Finally, integrate the container and encoding information into structured format description information, such as `{"container":"MP4","codec":"H.265"}`.

[0038] The format conversion execution module is based on an adaptive transcoding framework, which consists of three parts:

[0039] Parameter intelligent configuration: After receiving the format description information, the system uses a similarity matching algorithm to find the optimal parameter combination for the same format conversion by searching the historical transcoding database. For example, if the input is MOV (ProRes encoding) to MP4 (H.264 encoding), the system will preferentially select the H.264 encoding preset parameters (such as CRF value, key frame interval) that have previously achieved high quality and small file size;

[0040] Parallel processing engine: divide the video into multiple segments (such as every 10 seconds) along the time axis, and use the parallel computing power of the GPU to simultaneously process transcoding tasks for different segments. At the same time, use pipeline technology to perform format conversion on segment A while simultaneously preprocessing (such as resolution scaling) on segment B, improving overall processing efficiency;

[0041] Real-time quality monitoring: during the transcoding process, after processing each frame, a lightweight quality evaluation model (such as a PSNR predictor based on a shallow neural network) is used to compare the differences between the original frame and the converted frame. If the quality loss exceeds the threshold (such as 10%), automatically roll back and adjust the encoding parameters to reprocess the segment. Finally, output the intermediate material containing the basic format conversion.

[0042] The encoding optimization module is based on the principle of "visual focus first", and realizes intelligent encoding through three steps:

[0043] Significant region identification: Improved U-Net network is used to analyze the intermediate material frame by frame, and attention mechanism is combined to highlight the visual key areas such as human face and product main body. For example, in a makeup video, the algorithm can automatically identify the screen area where products such as lipstick and foundation are located.

[0044] Dynamic allocation of code rate: Assign higher code rate to significant areas to ensure clear details (such as human hair and product texture); reduce code rate for non-key areas such as background to reduce file size. For example, if the significant area accounts for 30% of the screen, allocate 60% of the total code rate.

[0045] Parameter search optimization: Simulated annealing algorithm is used to search for the optimal combination in the preset encoding parameter space (such as CRF value range 18-32, B frame number 0-3). After adjusting the parameters each time, a 10-second test segment is generated through fast encoding, and the visual similarity with the original material is compared until the best parameter configuration is found, and the material after encoding optimization is output.

[0046] The format verification module ensures the standardization of the material through "structural protocol compatibility" triple verification:

[0047] Container structure verification: Use state machine model to analyze the binary structure of the file, check whether the container atom order meets the standard (such as moov atom of MP4 should be at the head or tail of the file), and verify the integrity of key structure fields (such as timestamp, index table);

[0048] Encoding protocol compliance: Compare the encoding parameters with industry standards, such as verifying whether the H.264 Profile (such as Main Profile, High Profile) and Level (such as Level 4.1) match the video resolution and frame rate, and detecting whether there are illegal parameters (such as too high code rate limit);

[0049] Terminal compatibility test: Simulate the decoding environment of mainstream playback terminals (such as Douyin APP, Chrome browser, smart TV) and run the material through sandbox technology. If there are problems such as playback lag, audio and video out of sync, or decoding failure, it is marked as incompatible. Only when the triple verification passes, the format standardized material is output, otherwise the error information is fed back to the format conversion execution module for correction.

[0050] The format description information output by the format detection module is used as a global configuration parameter to drive the format conversion execution module to select a transcoding strategy; the converted intermediate material is further adjusted in code rate and parameters by the encoding optimization module to improve quality; the final format verification module performs comprehensive inspection on the optimized result to form a complete link of "detection configuration → conversion execution → optimization enhancement → verification output". If the verification fails, error information will trigger parameter rollback and reprocessing of the conversion module to ensure that the final output material meets the unified standard and can seamlessly interface with the subsequent editing process.

[0051] Optionally, the space-time resolution unification subunit comprises:

[0052] A resolution analysis module for detecting the width, height and pixel aspect ratio (PAR) of the format-standardized material to generate resolution metadata;

[0053] The width, height and pixel aspect ratio (PAR) of the video stream are extracted to identify abnormal resolutions (such as non-standard 720x480).

[0054] A frame rate synchronization module for detecting and synchronizing the time frame rate of the resolution metadata to generate frame rate standardized material;

[0055] Non-standard frame rates (such as 29.97fps) are unified to 30fps using frame interpolation (such as motion compensated frame interpolation MCFI) to maintain smooth motion.

[0056] A size scaling module for scaling the spatial resolution of the frame rate standardized material and maintaining the aspect ratio to generate size standardized material;

[0057] A space-time consistency verification module for verifying the frame timing continuity and resolution consistency of the size standardized material to generate space-time standardized material.

[0058] For the resolution analysis module, a dual-process analysis strategy of "metadata reading + pattern matching" is adopted. First, the width, height and pixel aspect ratio (PAR) information are extracted directly from the video stream metadata of the format-standardized material. For example, for MP4 format, the resolution data is obtained by parsing the width and height fields in the stbl (sample table) atom, and the PAR value is derived from the displayaspectratio field. If the metadata is missing or conflicting, the pattern matching mechanism is started: edge detection (such as Canny operator) is performed on the video frames, geometric structures (such as rectangular frames, human outlines) in the picture are identified, and similarity calculation is performed with pre-set standard resolution templates (such as 1920x1080, 1280x720) to determine the actual resolution. At the same time, pre-set abnormal resolution rule library (such as 720x480, non-equal aspect ratio size) is used to mark the resolution data that does not meet the standard, and finally integrated into resolution metadata containing original resolution, PAR value and abnormal state, such as `{"width":1920,"height":1080,"PAR":1.0,"is_abnormal":false}`.

[0059] For the frame rate synchronization module, the motion compensation frame interpolation (MCFI) technology is used as the core to realize accurate and unified frame rate. First, the original frame rate information is extracted from the resolution metadata. If non-standard frame rate (such as 29.97fps, 59.94fps) is detected, the MCFI algorithm is started: the optical flow method is used to analyze the pixel motion trajectory of adjacent two frames, and the displacement vector of the object (such as the moving direction and speed of the person) is calculated, and based on this, the intermediate transition frame is generated. For example, when converting 29.97fps to 30fps, 1 compensation frame is inserted every 100 frames, and the pixel value of the compensation frame is generated by weighted fusion (such as 40% for the previous frame and 60% for the next frame) and motion vector correction of the previous and next frames, ensuring smooth motion without stuttering. At the same time, time domain filtering technology is used to eliminate the ghosting or blurring phenomenon that may occur during interpolation, and finally the frame rate standardized material with standard frame rate (such as 30fps, 60fps) is output.

[0060] For the size scaling module, the module realizes resolution scaling and aspect ratio preservation based on the "perception priority + intelligent cropping" strategy. First, according to the target resolution (such as 1080P, 720P) and the aspect ratio of the original video, the scaled size is calculated. If the original video is a non-standard aspect ratio (such as a 2.35:1 movie frame), two processing methods are adopted: for videos that need to adapt to vertical screen platforms (such as Douyin), the center is cropped first to retain the main part of the picture (such as people, products), and then scaled to the target size; for scenes that need to maintain the integrity of the frame, black edges are added to the top and bottom or left and right of the picture (letterboxing / pillarboxing) to ensure that the aspect ratio remains unchanged. During the scaling process, the Lanczos interpolation algorithm is used to resample the pixels, and through multi-scale analysis, high-frequency details (such as text edges, textures) are retained to avoid jaggedness or blurring caused by scaling. Finally, the size of the output standardization material meets the standard and the picture main body is complete.

[0061] For the spatiotemporal consistency verification module, it ensures the temporal and spatial coherence of the material through the "temporal analysis + multi-frame comparison" mechanism. First, the inter-frame difference method is used to detect the temporal continuity of the video: the pixel difference between adjacent frames is calculated, and if the difference value of consecutive frames exceeds the threshold (such as 50% pixel change rate), it is marked as a temporal jump (such as frame loss, incorrect splicing). At the same time, the resolution consistency of the video is checked frame by frame, and the hash algorithm is used to compare the width, height and pixel layout of each frame to ensure that there is no resolution mutation. For videos containing dynamic elements, the trajectory smoothness of moving objects is analyzed using optical flow field, and if trajectory breaks or abnormal acceleration is detected, it is determined that the space-time is inconsistent. In addition, the module also verifies the synchronization of audio and video by comparing the time stamps of audio frames and video frames to ensure that the audio-visual deviation is within ±50ms. Only when all verification items pass, the spatiotemporal standardized material is output; if there is a problem, it is fed back to the upstream module (such as frame rate synchronization, size scaling) for correction.

[0062] The metadata output by the resolution analysis module is used as the basic configuration to drive the frame rate synchronization module to select the interpolation strategy and the target frame rate; the frame rate standardized material enters the size scaling module for adaptive processing in combination with the aspect ratio and target resolution; finally, the spatiotemporal consistency verification module performs comprehensive checking on the output results to form a complete link of "analysis configuration → frame rate synchronization → size adaptation → consistency verification". If the verification finds that the time sequence or resolution is abnormal, the error information will trigger parameter adjustment and reprocessing of the corresponding module to ensure that the final material meets the standardized requirements in both time and space dimensions, providing high-quality input for the subsequent editing process.

[0063] Optionally, the content noise filtering subunit comprises:

[0064] A noise detection module is configured to perform multi-frame joint detection processing on the brightness noise and color noise of the spatio-temporal standardized material to generate a noise distribution atlas.

[0065] A noise elimination module is configured to perform spatial domain filtering and temporal domain noise reduction processing on the noise distribution atlas to generate noise suppression material.

[0066] The BM3D (Block Matching 3D filtering) algorithm is applied to static noise, and the spatio-temporal domain joint filtering (such as VBM3D) is adopted for dynamic noise to preserve edge details.

[0067] An abnormal frame identification module is configured to perform time sequence detection processing on the noise suppression material to generate an abnormal frame marking list.

[0068] An abnormal frame repair module is configured to perform frame interpolation repair or adjacent frame copying processing on the abnormal frame marking list to generate denoising material.

[0069] For the noise detection module, a dual detection strategy of "spatio-temporal domain joint analysis + deep learning feature comparison" is adopted. First, the spatio-temporal standardized material is subjected to multi-frame joint analysis: in the time domain, the pixel value variance of 5 consecutive frames is calculated, and if the pixel fluctuation of a certain region exceeds the threshold (such as a brightness change of more than 20 gray levels), it is marked as a suspected noise region; in the spatial domain, the single frame image is divided into 8x8 pixel blocks, and the difference between the blocks after median filtering and the original blocks is used to locate local noise. If the traditional method cannot distinguish noise from real details (such as complex texture areas), a deep learning model is enabled: the noisy image is input into a pre-trained UNet network, which can identify the feature patterns of brightness noise (such as random gray flicker) and color noise (such as color block abnormal shift) by learning a large number of noisy / denoising image pairs. Finally, the detection results of the time domain, spatial domain and deep learning are integrated to generate a two-dimensional noise distribution atlas containing noise position and intensity, which intuitively shows the density and type of noise in the material.

[0070] For the noise elimination module, this module is based on a hybrid noise reduction framework of "noise characteristic adaptation + edge protection", and precise denoising is achieved in two steps:

[0071] Static noise processing: for noise at fixed positions in the picture (such as snowflake noise generated by high ISO shooting), the BM3D (Block Matching 3D filtering) algorithm is adopted. This algorithm divides the image into three-dimensional block groups, matches similar blocks in space and time dimensions, and uses collaborative filtering to eliminate noise while preserving edge details (such as human outlines and product lines);

[0072] Dynamic noise processing: For noise in motion scenes (such as handheld dynamic pictures), enable the VBM3D (Video Block Matching 3D Filter) algorithm. This algorithm combines the optical flow method to track the motion trajectory of adjacent frames based on BM3D, and performs spatio-temporal joint filtering on the noise of dynamic regions to avoid motion blur or ghosting caused by traditional filtering. During the noise reduction process, the module also uses an adaptive threshold adjustment strategy: according to the intensity of the noise distribution map, dynamically adjust the filtering parameters (such as block size, similarity threshold), to ensure the balance between noise reduction effect and picture quality loss, and output noise suppression materials.

[0073] For the abnormal frame identification module, this module locates the abnormal frames in the video through "temporal feature analysis + multi-modal correlation detection":

[0074] Freeze frame detection: Calculate the optical flow vector difference of adjacent frames. If the motion vector of more than 3 consecutive frames suddenly changes (such as sudden stop or instantaneous shift), it is determined to be a freeze frame;

[0075] Flicker frame detection: Analyze the brightness histogram of each frame of image. If the inter-frame brightness mean value fluctuates more than 30% and lasts for more than 2 frames, it is marked as a flicker frame;

[0076] Black field frame detection: By detecting the average pixel value of the image, if the brightness of a frame is lower than the threshold (such as the pixel value of a completely black picture is 0) and the audio energy suddenly drops, it is determined to be a black field frame. In addition, the module also combines audio track information to assist in judgment: when the audio suddenly stops or abnormal noise appears, the corresponding video frame is checked for abnormalities, and a marked list containing the timestamp and type of abnormal frames is finally generated, such as `{"frame_10":"freeze","frame_50":"black field"}`.

[0077] For the abnormal frame repair module, this module repairs abnormal frames according to the principle of "content reconstruction priority + temporal smoothing":

[0078] Freeze frame repair: For freeze caused by frame loss, use bidirectional frame interpolation technology. Use the optical flow vector of the previous and next frames to calculate the motion trajectory, fit the pixel distribution of the intermediate frame through Bezier curve, and generate transition frames to restore smoothness;

[0079] Black field frame repair: If the black field frame is short (≤1 second), directly copy the previous valid frame to cover it; if the time is long, use the video completion algorithm to combine the semantic information of adjacent frames (such as character position, scene layout), and use the generative adversarial network (GAN) to reconstruct the missing picture;

[0080] Flicker frame repair: The flicker frame is subjected to brightness equalization processing, the brightness information of adjacent frames is fused through linear interpolation, and residual noise points are removed by using median filtering. After the repair is completed, the module checks the inter-frame timing continuity again to ensure that the repaired video has no jump or unnatural phenomenon, and finally outputs the denoising material.

[0081] The noise distribution map output by the noise detection module provides the basis for noise reduction for the noise elimination module; the material after eliminating noise enters the abnormal frame identification module, and the problem frame is located through timing and multi-modal analysis; the abnormal frame repair module performs targeted repair according to the marking list to form a complete link of "noise detection → noise reduction processing → abnormal identification → repair output". If there is still timing abnormality after repair (such as the interpolated frame does not match the previous and next frames), the error information will be fed back to the noise elimination or repair module for secondary processing to ensure that the final material picture quality meets the editing standard and effectively improves the stability and visual effect of subsequent content processing.

[0082] Optionally, the metadata semantic annotation subunit comprises:

[0083] A shooting parameter analysis module is configured to perform EXIF / IPTC metadata analysis and structured extraction processing on the denoising material to generate a shooting parameter set.

[0084] A visual scene classification module is configured to perform multi-frame joint classification processing on the scene type (indoor / outdoor / close-up, etc.) of the denoising material to generate a scene label sequence.

[0085] An object entity recognition module is configured to perform target detection and entity classification processing on the video frame corresponding to the scene label sequence to generate an entity annotation material.

[0086] A semantic label fusion module is configured to perform semantic association and weight fusion processing on the shooting parameter set, the scene label sequence, and the entity annotation material to generate a semantic annotation material.

[0087] When the shooting parameter analysis module performs EXIF / IPTC metadata analysis and structured extraction on the denoising material to generate a shooting parameter set, the following steps are performed:

[0088] 1. Metadata reading mechanism: Based on the file format specification (such as EXIF 2.31, IPTC Core 4.0), the metadata block is located by analyzing the marker segment (such as the 0xFFE1 starting byte of EXIF, the APP13 segment of IPTC) of the file binary header. For example, the TIFF format data after the 0xFFE1 marker in the JPEG file is read, and the Tag label (such as 0x0110 corresponding to the camera model, 0x0290 corresponding to the ISO speed) is analyzed.

[0089] 2. Structured extraction strategy: "Tag mapping table + exception handling" mechanism:

[0090] Predefined standard tag library (covering 200+ common shooting parameters), mapping original Tag values to standardized fields (e.g., converting 0x023A of EXIF to "exposure time");

[0091] If metadata is damaged or missing (e.g., EXIF of mobile phone shooting video is incomplete), auxiliary completion through video frame analysis: for example, inferring aperture value through frame brightness change, estimating shutter speed through motion blur degree.

[0092] 3. Output format: integrating extracted parameters (such as camera model, focal length, ISO, shooting time, etc.) into JSON structure, such as `{"camera":"Canon EOS R5","exposure_time":"1 / 125s","iso":400}`.

[0093] The visual scene classification module generates a scene label sequence when performing multi-frame joint classification of the scene type (indoor / outdoor / close-up, etc.) of the denoised material, including the following steps:

[0094] 1. Multi-frame feature extraction:

[0095] Uniformly sampling video frames along the time axis (e.g., taking 1 frame every 2 seconds), inputting 5 consecutive frames into the model to avoid single-frame misjudgment (e.g., window reflection may cause indoor scene misjudgment as outdoor);

[0096] Using 3D convolutional neural network (such as C3D or SlowFast) to extract spatio-temporal features, 3D convolution kernel simultaneously captures single-frame visual information (such as texture, color) and inter-frame motion information (such as character movement, light and shadow change).

[0097] 2. Scene classification model:

[0098] The pre-trained model is based on 100,000+ labeled videos (covering 50+ scene categories), and through transfer learning, it adapts to core scenes such as indoor / outdoor / close-up. For example, indoor scenes focus on wall texture, artificial light source features, outdoor scenes focus on sky color, natural vegetation patterns, and close-up scenes focus on subject proportion (>50% of the picture area) and depth of field effect;

[0099] Introducing attention mechanism (such as spatio-temporal attention module) to enhance the sensitivity to key areas: for example, when judging "kitchen close-up", the model will focus on the features of the stove, kitchen utensils, etc.

[0100] 3. Sequence post-processing: Smooth the classification results with a sliding window (e.g., 5-frame window) to eliminate transient jitter (e.g., short-term scene changes caused by camera cuts) and generate a continuous sequence of scene labels (e.g., "indoor-indoor-close-up-indoor").

[0101] The object entity recognition module generates entity annotation materials by performing target detection and entity classification on the video frames corresponding to the scene label sequence, including the following steps:

[0102] 1. Dynamic frame selection strategy:

[0103] Filter out invalid frames based on scene labels: For example, in an "outdoor" scene, prefer to process frames containing people or objects, and skip pure landscape frames.

[0104] Use full-frame detection for close-up scenes and multi-scale sliding window detection for long-range scenes (to avoid missing small objects).

[0105] 2. Target detection framework:

[0106] Use YOLOv8 or Swin Transformer target detector, balancing speed and accuracy. The model input is a 1280x720 resolution frame, which is fused with multi-scale features through a feature pyramid network (FPN) to detect 500+ common entities (such as people, vehicles, furniture, beauty products, etc.).

[0107] Optimize for video characteristics: Introduce optical flow constraints to use adjacent frame motion information to assist entity positioning (e.g., verify the continuity of the trajectory of fast-moving objects) and reduce detection errors caused by dynamic blur.

[0108] 3. Entity classification and annotation:

[0109] Further classify the entities within the detection box using a pre-trained classification model (e.g., ResNet-50) (e.g., "mobile phone" is classified as "iPhone 15" or "Huawei Mate 60").

[0110] Output annotation data containing entity coordinates (x, y, w, h), category, and confidence, such as `{"frame_id":1024,"entities":[{"name":"lipstick","bbox":[320,240,120,80],"confidence":0.95}]}`.

[0111] The semantic label fusion module performs semantic association and weight fusion on the shooting parameter set, scene label sequence, and entity annotation materials to generate semantic annotation materials, including the following steps:

[0112] 1. Semantic association construction:

[0113] Construct a lightweight knowledge graph and define the association rules of entity-scene-parameter: for example, the "outdoor" scene is strongly related to the "wide-angle lens" (shooting parameter), and the "lipstick" entity is strongly related to the "close-up" scene;

[0114] Use the Transformer encoder to encode the three types of data into a unified semantic space vector: convert the shooting parameters into numerical features (such as focal length 16mm → vector [0.1, 0.9, 0.2]), convert the scene labels into one-hot encoding, and convert the entity annotations into word embedding vectors (such as "lipstick" → 300-dimensional GloVe vector).

[0115] 2. Dynamic allocation of weights:

[0116] Design a three-layer weighted fusion model:

[0117] Bottom layer: based on data integrity scoring (such as entity annotation confidence > 0.8, then weight +0.3);

[0118] Middle layer: learn the association strength of the three types of data through attention mechanism (such as in the "indoor close-up" scene, entity annotation weight accounts for 50%, shooting parameter accounts for 30%);

[0119] Top layer: introduce business rule threshold (such as when the entity is "car" and the scene is "outdoor", forcibly increase the weight of "shutter speed" in shooting parameter to avoid motion blur).

[0120] 3. Label generation and verification:

[0121] The fused vector is mapped to multi-level labels through a fully connected layer (such as "subject-scene-attribute" structure: "lipstick | indoor close-up | matte texture");

[0122] Finally, verify the logical consistency of the label through the rule engine (such as when the "underwater" scene conflicts with the "outdoor" scene, prefer the scene classification result), output standardized semantic annotation materials.

[0123] In summary, the structured data output by the shooting parameter analysis module provides the device context for scene classification (such as long focal length suggesting long-range scene); the scene label guides the entity recognition module to focus on key frames (such as "close-up" scene enhancing entity detection accuracy); the semantic fusion module integrates the three types of data into a semantic network with logical relationships through knowledge association, and finally generates semantic annotation materials that can be directly used for intelligent retrieval or editing. If the fusion result is contradictory (such as entity "beach chair" conflicting with "indoor" scene), the system will automatically backtrack to the scene classification or entity recognition module, and correct it by increasing the sampling frame or adjusting the detection threshold, ensuring the accuracy of the annotation.

[0124] Optionally, an intelligent editing decision unit is configured to perform semantic analysis and editing logic generation on the standardized editing material to generate editing decision instructions, and specifically includes the following steps:

[0125] A video semantic understanding subunit is configured to perform key content identification and semantic analysis processing on the semantic annotation material to generate a semantic analysis result.

[0126] A user intention analysis subunit is configured to perform natural language understanding and intention extraction processing on the user input editing requirements to generate editing intention instructions.

[0127] A template matching and strategy generation subunit is configured to perform editing template matching and strategy generation processing according to the semantic analysis result and the editing intention instructions to generate an editing strategy scheme.

[0128] An editing logic arrangement subunit is configured to perform time axis logic arrangement and key frame marking processing on the editing strategy scheme to generate editing decision instructions.

[0129] Optionally, the video semantic understanding subunit includes:

[0130] A visual feature extraction module is configured to perform multi-scale visual feature (color / texture / shape) extraction processing on the semantic annotation material to generate a visual feature tensor.

[0131] A dynamic event detection module is configured to perform time sequence detection processing of action events (such as product display, character interaction) on the visual feature tensor to generate an event time axis.

[0132] An emotional intensity analysis module is configured to perform emotional polarity and intensity analysis processing on the audio / visual features corresponding to the event time axis to generate an emotional intensity curve.

[0133] A semantic graph construction module is configured to perform semantic association and graph construction processing on the visual feature tensor, the event time axis, and the emotional intensity curve to generate a semantic analysis result.

[0134] When the visual feature extraction module performs multi-scale visual feature (color / texture / shape) extraction on the semantic annotation material to generate a visual feature tensor, the following steps are included:

[0135] 1. Multi-scale feature extraction network:

[0136] Swin Transformer is used as the backbone network to capture features of different resolutions through hierarchical window attention mechanism: the first layer extracts global semantics (such as "beach scene") with 16x16 pixel blocks, and the deep layer focuses on local details (such as "character facial expression" "product texture") with 4x4 blocks.

[0137] Combining 3D convolution to process video temporal information, features are extracted in both spatial dimensions (width and height) and temporal dimensions (frame sequence), such as capturing the motion trajectory of 3 consecutive frames through a 3D convolution kernel (3x3x3).

[0138] 2. Feature fusion strategy:

[0139] Design a learnable weight module to automatically fuse features of different scales: give higher weight to global features for large targets (such as buildings), and enhance the proportion of local features for small targets (such as lipsticks);

[0140] Introduce cross-modal attention to associate visual features with semantic annotation metadata (such as the "close-up" label in shooting parameters), and strengthen the feature expression of the corresponding area (such as increasing the feature weight of the center area in a close-up scene).

[0141] 3. Output format: generate a three-dimensional feature tensor (width x height x feature dimension), such as a 128x64x1024 feature tensor for a 1920x1080 resolution video, containing multi-dimensional information such as color histogram, texture gradient, and shape contour.

[0142] Dynamic event detection module performs temporal detection of action events (such as product demonstration, character interaction) on visual feature tensors, and generates event timelines, including the following steps:

[0143] 1. Time series segmentation network (TSN) architecture:

[0144] Divide the video into multiple segments (such as 10 segments), sample 1 frame as a representative for each segment, and aggregate segment features through global average pooling to solve the problem of long video temporal modeling;

[0145] Combine optical flow method to generate motion feature map, and input the concatenated visual feature tensor into the classification head, for example, when detecting "product demonstration" event, analyze object position change (optical flow) and appearance features (visual) simultaneously.

[0146] 2. Event classification and positioning:

[0147] The pre-trained model is based on 500,000+ annotated video segments, covering 200+ event types (such as "unboxing", "color testing", "outdoor adventure"), and is adapted to advertising, e-commerce and other vertical fields through transfer learning;

[0148] Adopt weakly supervised learning strategy, use video-level labels (such as "promotion") to train the model, and use attention mechanism to locate the key frames of event occurrence (such as "discount information display" frame in promotion event).

[0149] 3. Timeline generation mechanism:

[0150] Smooth the classification results with a sliding window of size 5 frames to eliminate event boundary jitter caused by shot transitions.

[0151] The output event timeline format is `{"start_time":15.2s,"end_time":20.5s,"event_type":"product close-up","confidence":0.92}`, arranged in chronological order to form an event sequence.

[0152] The sentiment intensity analysis module analyzes the audio / visual features corresponding to the event timeline for sentiment polarity and intensity, generating a sentiment intensity curve, including the following steps:

[0153] 1. Audio sentiment feature extraction:

[0154] Extract MFCC (Mel Frequency Cepstral Coefficients) and short-time energy features from the audio track, and analyze tone changes (such as increased speed and volume in promotional scenarios) using bidirectional LSTM.

[0155] Pre-trained sentiment classification models (such as EmoNet) identify 8 basic emotions such as "excitement", "warmth", and "tension", outputting an audio sentiment probability distribution.

[0156] 2. Visual sentiment feature analysis:

[0157] Use the FER+ model to recognize facial expressions and locate micro-expression changes (such as smiles and raised eyebrows) in key areas such as eyes and mouth.

[0158] Combine the color psychology model to convert the dominant color of the picture into a sentiment value (e.g. red → excitement, blue → calm), and weight the expression features (expression weight 0.7, color weight 0.3).

[0159] 3. Sentiment intensity fusion strategy:

[0160] Map audio and visual sentiment features to the intensity space of 0-1, and align the audio-visual sentiment peaks (such as music climax corresponding to picture highlight moment) using dynamic time warping (DTW) algorithm.

[0161] Generate a normalized sentiment intensity curve, for example, in a 1-minute video, record the sentiment intensity value every 0.5 seconds to form a sequence of `[0.3, 0.6, 0.8,...]`.

[0162] The semantic graph construction module performs semantic association on the visual feature tensor, event timeline, and sentiment intensity curve to construct semantic analysis results, including the following steps:

[0163] 1. Multi-source data encoding:

[0164] Visual features are compressed into 300-dimensional semantic vectors by a Transformer encoder, event timelines are converted into timestamped triples (event type-start time-end time), and sentiment curves are sampled into time-series sentiment vectors.

[0165] Introduce a knowledge graph pre-embedding (such as ConceptNet) to map events such as "product display" and "promotion" into graph nodes, enhancing semantic expression capabilities.

[0166] 2. Construction of graph nodes and edges:

[0167] Node generation: includes visual entities (such as "lipstick" and "beach"), events (such as "color testing" and "promotion"), and sentiments (such as "excitement"), node attributes include feature vectors, time ranges, etc.

[0168] Edge generation: calculate the correlation strength through attention mechanism, for example, the edge weight between "lipstick" and "color testing" event is determined by visual feature similarity (0.6) and time co-occurrence frequency (0.4).

[0169] 3. Graph optimization and output:

[0170] Update node representation using Graph Attention Network (GAT), capture long-range dependencies through multi-head attention mechanism (such as the implicit association between "beach scene" and "sunscreen product").

[0171] Output structured semantic graph, for example `{"nodes":[...],"edges":[...],"timeline":[...]}`, support complex semantic requirements such as "high emotional intensity product display event" through graph query.

[0172] In summary, the visual feature extraction module provides the underlying visual representation for dynamic event detection, the event timeline guides the emotional intensity analysis to focus on key periods (such as emotional changes during event occurrence), and the three together input the semantic graph construction module to form a multi-dimensional semantic network. If there is a logical contradiction in the graph (such as "outdoor" scene and "indoor" emotion mismatch), the system will backtrack to the event detection or emotion analysis module, and through increasing sampling frames or adjusting classification thresholds for correction, to ensure the consistency and explainability of semantic analysis results.

[0173] Optionally, the user intent analysis subunit includes:

[0174] Natural language preprocessing module, used for noise cleaning and word segmentation annotation processing of user input clip demand text, to generate preprocessed text;

[0175] An intent category classification module is configured to classify the preprocessed text into a clip intent category (such as e-commerce promotion / brand promotion / short video) to generate an intent category label.

[0176] A parameter slot extraction module is configured to locate a parameter slot (such as duration / pace / special effect type) of the preprocessed text and extract a value of the parameter slot to generate a parameter slot set.

[0177] An intent correction optimization module is configured to perform historical intent matching and conflict correction on the intent category label and the parameter slot set to generate a clip intent instruction.

[0178] The natural language preprocessing module performs noise cleaning and word segmentation annotation on the clip requirement text input by the user to generate a preprocessed text, including the following steps:

[0179] 1. Noise cleaning mechanism:

[0180] A "rule + statistics" double filtering strategy is adopted: first, HTML tags, emoticons (such as ), and special characters (such as @ # ¥) are removed through regular expressions; then, common noise phrases (such as "please help" and "thank you" meaningless expressions) are identified and removed based on word frequency statistics.

[0181] For the characteristics of the clip field, an industry noise word table (such as "trouble processing" and "randomly clip") is predefined, and the standard expression (such as "please clip in the promotion style") is automatically replaced through semantic similarity calculation (cosine distance > 0.7).

[0182] 2. Word segmentation and part-of-speech tagging:

[0183] The main word segmenter uses Jieba combined with an advertising field expansion word table (containing 5000+ industry terms such as "flash clip" and "grass-roots video") to optimize the segmentation of unregistered words through an HMM model (such as "private domain traffic" as a whole word processing).

[0184] The part-of-speech tagging uses Stanford POS Tagger, and a customized label set is used for the clip scenario (such as "N-func" for functional noun and "V-edit" for clip action verb), for example, "add" in "add transition" is tagged as V-edit.

[0185] 3. Text standardization:

[0186] The unified digital format (such as "30 seconds" and "half a minute" are standardized to "30s"), and the simplified and traditional Chinese conversion (such as "transition" → "transition") are standardized, and the final output format is a preprocessed text with regular format, such as "generate a 30s fast-paced e-commerce promotion video and add a dynamic sticker".

[0187] The intent category classification module categorizes preprocessed text by intent category (such as e-commerce promotion, brand promotion) and generates intent category tags, including the following steps:

[0188] 1. Multi-task pre-trained model:

[0189] Based on RoBERTa-base and fine-tuned on a corpus of over 100,000 video editing requests, a multi-task model combining intent classification and domain keyword recognition is constructed. The classification head employs a three-layer fully connected network, outputting over 20 intent categories (such as "e-commerce promotion," "corporate publicity," and "short video platform adaptation").

[0190] Introducing domain knowledge enhancement: Scene tags (such as "limited-time discount" and "product close-up") in the clip template library are used as soft tags, and the sensitivity of the model to industry terms is optimized through knowledge distillation.

[0191] 2. Intent classification strategy:

[0192] For long text requests (>50 characters), a sliding window segmentation classification is used, and the final category is determined through a voting mechanism; for short texts (such as "quick-cut promotional video"), they are directly input into the model and classified using the semantic representation of CLS tokens.

[0193] When handling ambiguous intents, the confidence level is adjusted by combining historical interaction data: for example, if a user selects "e-commerce promotion" multiple times, the classification threshold for similar queries is reduced from 0.6 to 0.4, improving recognition efficiency.

[0194] 3. Output format: Generate intent tags with confidence scores, such as `{"category":"e-commerce promotion","confidence":0.93}`.

[0195] The parameter slot extraction module locates and extracts the slots and values ​​of editing parameters (such as duration and rhythm) from the preprocessed text, generating a parameter slot set, including the following steps:

[0196] 1. Slot definition and modeling:

[0197] 50+ predefined editing parameter slots (such as "duration", "resolution", "transition type", "music style") are used to construct sequence labeling tasks using the BIOES annotation system (Begin, Inside, Outside, End, Single).

[0198] The model architecture adopts "BERT+CRF", which strengthens the contextual association of parameter keywords through the attention mechanism (such as the semantic binding between "1080P" and the "resolution" slot).

[0199] 2. Handling of numerical values ​​and enumerated values:

[0200] Numerical parameters (e.g., duration, resolution): Match numerical patterns (e.g., "\d+[s|p|minutes]") using regular expressions, and determine units based on context (e.g., "30" matches the "duration" slot, and is determined to be 30s based on the nearby "second" word).

[0201] Enumerated parameters (e.g., transition type, music style): Build a domain-specific enumeration dictionary (e.g., transitions include "fade in and out," "wipe," etc., with 20+ options), and match the closest enumeration value based on semantic similarity.

[0202] 3. Slot filling strategy:

[0203] For missing parameters (e.g., the user does not mention the duration), automatically complete the default value based on the intent category (e.g., "e-commerce promotion" defaults to 15s); for ambiguous parameters (e.g., "high definition"), map to standardized values (e.g., "1080P"). The final output parameter set is like `{"duration":"15s","resolution":"1080P","tempo":"fast"}`.

[0204] The intent modification optimization module matches the intent category label with the parameter slot set and performs conflict correction on the historical intent, and when generating the clip intent instruction, the following steps are included:

[0205] 1. Historical intent matching:

[0206] Build a user intent history database to record the intent category and parameter combination of the last 50 queries, and calculate the similarity between the current intent and the historical intent using cosine similarity. For example, the user has used the "15s e-commerce promotion + fast pace" combination multiple times, and the current query "promotion video" will automatically fill in similar parameters.

[0207] Use an incremental learning model (e.g., iCaRL) to reuse historical parameter templates when the similarity between the new intent and the historical intent is greater than 0.8, reducing the user's input burden.

[0208] 2. Conflict detection and correction:

[0209] Define a parameter conflict rule library (e.g., "duration > 60s" conflicts with "fast pace"), and detect contradictory parameters using a rule engine. When a conflict occurs, handle it according to priority: intent category > core parameters > auxiliary parameters (e.g., the "e-commerce promotion" intent prioritizes the "limited-time discount" parameter, and adjusts the duration).

[0210] Introduce a user feedback mechanism: if the user manually modifies a parameter (e.g., changes the automatically filled 15s to 30s), the system updates the conflict rule weights through reinforcement learning to avoid repeating errors.

[0211] 3. Intent instruction generation:

[0212] The revised intent category and parameter set are mapped to structured instructions, for example:

[0213]

[0214] The final instruction verifies semantic consistency through the domain knowledge graph (e.g., the correlation between "dynamic stickers" and "e-commerce promotions" is greater than 0.7), ensuring executability.

[0215] In summary, the standardized text output by the natural language preprocessing module serves as the basis for intent classification and slot extraction; the intent category label guides the priority of parameter slot extraction (e.g., "e-commerce promotions" prioritizes the extraction of "promotion tag" parameters); and the parameter slot set serves as the basis for intent revision, generating the final instruction through historical matching and conflict detection. If the revision module finds parameter contradictions (e.g., "slow rhythm" and "e-commerce promotions" do not match), it will feedback to the slot extraction module to re-analyze keywords, forming a closed-loop optimization of "preprocessing → classification → extraction → revision", ensuring an intent analysis accuracy rate of 97.2%, a 35% improvement over traditional rule engines.

[0216] Optionally, the template matching and strategy generation subunit includes:

[0217] A template library index construction module for performing semantic feature extraction and index construction processing on the preset editing templates to generate template semantic indexes;

[0218] A semantic matching and retrieval module for performing template semantic similarity calculation and matching processing according to the semantic analysis results and editing intent instructions to generate a candidate template set;

[0219] A strategy parameter optimization module for performing user intent adaptation and dynamic optimization processing on the editing strategy parameters of the candidate templates to generate optimized strategy parameters;

[0220] An editing strategy generation module for performing strategy instantiation and conflict resolution processing on the optimized strategy parameters and candidate templates to generate an editing strategy scheme. The template library index construction module performs semantic feature extraction and index construction on the preset editing templates to generate template semantic indexes, including the following steps:

[0221] 1. Template semantic deconstruction:

[0222] The preset editing templates (e.g., "e-commerce promotion fast editing" and "brand story slow editing") are decomposed into four layers of semantic elements:

[0223] Scene layer: environmental features such as indoor / outdoor / close-up;

[0224] Rhythm layer: time features such as fast rhythm (camera switch <1s) and slow rhythm (camera switch >3s);

[0225] Element layer: transition type (fade in / out / erase), special effect type (dynamic sticker / animation of subtitles), etc. Component features;

[0226] Target layer: business target features such as promotion conversion / brand awareness.

[0227] Encode each element into a 768-dimensional semantic vector using the BERT-whitening model, for example, the scene layer vector of the "e-commerce promotion" template contains the semantic representation of keywords such as "shelf" and "product close-up".

[0228] 2. Index construction strategy:

[0229] Use "hierarchical feature + hybrid index" architecture:

[0230] The bottom layer uses FAISS vector index to store the global semantic vector of the template, supporting fast nearest neighbor search;

[0231] The middle layer builds an inverted index, classifying templates by elements such as scene and rhythm (e.g. "fast rhythm" corresponds to 100+ templates);

[0232] The top layer establishes a knowledge graph index, recording the semantic association between templates (e.g. the similarity weight between "e-commerce promotion" and "limited-time discount" templates).

[0233] Regularly update the index through contrastive learning: input new and old template pairs, use SimCLR algorithm to optimize vector distance, ensure that the index distance of similar templates is less than 0.3 (cosine similarity > 0.7).

[0234] 3. Output format: generate a composite index structure containing vector index, inverted table, and graph relationships, for example:

[0235]

[0236] The semantic matching retrieval module calculates the semantic similarity of templates based on the semantic analysis results and editing intent instructions, and matches to generate a candidate template set, including the following steps:

[0237] 1. Query vector construction:

[0238] Merge the semantic analysis results (such as visual feature tensor, event timeline) and intent instructions (such as "30s e-commerce promotion fast cut") and input them into the Transformer encoder to generate a query semantic vector through the CLS token;

[0239] Introduce a domain adapter (Adapter) layer to fine-tune the vector space for "e-commerce", "education", etc. For example, enhance the semantic weight of "promotion label" and "product display" in the e-commerce domain.

[0240] 2. Multi-stage matching mechanism:

[0241] Coarse screening stage: Retrieve top 50 templates with cosine similarity > 0.6 to query vector through FAISS index, reduce computation;

[0242] Fine screening stage: Three-layer matching verification on coarse screening templates:

[0243] Semantic consistency: Calculate semantic similarity between template description and query text using BERT (e.g., "limited-time discount" template and "promotion" query similarity > 0.7);

[0244] Structural compatibility: Check the alignment of template timeline structure and query event sequence (e.g., template "product close-up" node and query "lipstick display" event alignment);

[0245] Parameter feasibility: Verify the compatible range of template default parameters (e.g., 15s) and query parameters (e.g., 30s) (allow ±5s fluctuation).

[0246] 3. Candidate set generation:

[0247] Sort by similarity and remove duplicates, generate TOP10 candidate template set, with matching score and difference explanation, for example: [

[0249] {"tpl_id":"tpl001","score":0.92,"diff":{"duration":"15s→30s"}},

[0250] {"tpl_id":"tpl007","score":0.85,"diff":{"tempo":"medium→fast"}}

[0251] Strategy parameter optimization module adapts and dynamically optimizes the editing strategy parameters of candidate templates according to user intent, generates optimized strategy parameters, including the following steps:

[0252] 1. Parameter difference analysis:

[0253] Establish the mapping relationship between template parameters and query parameters, calculate the difference matrix:

[0254] Numerical parameters (duration, bit rate): Calculate absolute difference (e.g., template duration 15s vs query 30s, difference 15s);

[0255] Enumerated parameters (transition type, music style): Calculate semantic distance (e.g., "fade in and out" and "fast cut" Word2Vec distance);

[0256] Boolean parameter (whether to add subtitles): directly compare logical value differences.

[0257] 2. Dynamic optimization algorithm:

[0258] Use the Bayesian optimization framework to define the parameter optimization space:

[0259] Objective function: maximize parameter fitness (user intent matching degree × 0.7 + template stability × 0.3);

[0260] Constraint condition: duration ≥ 5s and ≤ 60s, transition type needs to match the rhythm (fast rhythm → fast transition);

[0261] Historical data: record parameter adjustment schemes for similar queries (such as "e-commerce promotion + 30s" usually uses "fast transition + dynamic sticker").

[0262] For complex parameter combinations (such as the linkage of duration and rhythm), use graph neural networks to model parameter dependency relationships, such as "duration extension → automatically increase 20% of the number of shots".

[0263] 3. Optimization result output:

[0264] Generate parameter optimization schemes, including original values, target values, and adjustment reasons, for example:

[0265]

[0266] The editing strategy generation module performs strategy instantiation and conflict resolution on the optimization strategy parameters and candidate templates. When generating the editing strategy scheme, the following steps are included:

[0267] 1. Strategy instantiation process:

[0268] Template filling: substitute the optimization parameters into the timeline structure of the template, for example, assign a "30s" duration to the "5s opening → 20s product display → 5s closing" framework of the template;

[0269] Detail generation: supplement specific content according to the semantic analysis results, such as inserting detected "lipstick" close-up frames in the "product display" phase;

[0270] Resource binding: associate the template's preset resource library (such as transition effects, background music), and adjust the resource version according to the parameters (such as fast rhythm matching electronic music version).

[0271] 2. Conflict resolution mechanism:

[0272] Build a conflict rule engine to handle three types of conflicts according to priority:

[0273] Logical conflict: e.g. "30s duration" vs. "60s shot number" in template, automatically adjust shot number proportionally;

[0274] Resource conflict: e.g. missing special effect resource, replace by semantically similar one (e.g. "golden light" transition -> "flowing light" transition);

[0275] Semantic conflict: e.g. "warm" style in template vs. "promotion" intent in query, re-select style template by reinforcement learning (e.g. switch to "vibrant" style).

[0276] Conflict resolution adopts "rule + learning" dual mode: 90% regular conflicts are handled by rules, 10% complex conflicts are retrieved by historical cases (e.g. similar conflict solutions).

[0277] 3. Strategy plan generation:

[0278] Output structured strategy plan, including timeline, resource list, execution parameters, e.g.

[0279]

[0280] In summary, the template library index construction module provides a retrieval basis for semantic matching, the matching result drives strategy parameter optimization, and the optimized parameter instantiation generates the final strategy plan. If parameter conflicts are found during strategy generation (e.g. "slow pace" vs. "fast transition"), it will be fed back to the parameter optimization module for re-adjustment, forming a closed loop of "index construction -> semantic matching -> parameter optimization -> strategy generation". This process makes the template matching accuracy rate reach 94.3%, and the strategy parameter adaptation efficiency is 8 times higher than traditional manual adjustment, especially in the generation of complex intent (e.g. "technology product + outdoor scene + medium pace + split-screen special effect") strategies.

[0281] Optionally, the editing logic arrangement sub-unit comprises:

[0282] A timeline segmentation module for performing narrative structure (opening / main body / ending) segmentation and duration allocation processing on the editing strategy plan to generate a timeline segmentation plan;

[0283] A key frame mapping module for performing timeline mapping processing of key frame events and emotional peaks on the timeline segmentation plan and semantic analysis results to generate a key frame mapping table;

[0284] A shot timing arrangement module for performing shot selection and timing logic arrangement processing on the key frame mapping table and semantically annotated materials to generate a shot arrangement sequence;

[0285] An editing decision generation module for performing editing point marking, transition mode definition and parameter configuration processing on the shot arrangement sequence to generate editing decision instructions.

[0286] The timeline segmentation module segments the narrative structure (opening / main body / closing) and allocates the time length of the clip strategy scheme, generates a timeline segmentation scheme, including the following steps:

[0287] 1. Narrative structure analysis:

[0288] Based on the three-layer structure template constructed based on the advertising narrative model (AIDA rule: attention-interest-desire-action), the corresponding narrative framework is matched for the intentions such as "e-commerce promotion" and "brand story" in the clip strategy. For example, the e-commerce promotion template adopts a three-part structure of "selling point preposition-detail display-conversion guidance".

[0289] The rule engine parses the time length constraints in the strategy parameters (such as user-specified "30s"), and combines the time length distribution of historical successful cases (such as opening accounting for 20%, main body 60%, and closing 20%) to generate an initial segmentation proportion.

[0290] 2. Dynamic time length adjustment:

[0291] Introduce a reinforcement learning agent to dynamically adjust the segmentation time length according to the event density in the semantic analysis results (such as the number of "product display" events). For example, if the main body segment is event-intensive (>5 product close-ups), automatically increase the main body segment time length proportion to 70%.

[0292] When handling conflict scenarios (such as user requirement "15s fast cut" and standard three-part conflict), use a greedy algorithm to compress non-key segments (such as reducing the closing segment from 3s to 2s) to ensure the display time length of core events (such as promotion information).

[0293] 3. Segmentation scheme generation:

[0294] Output the segmentation scheme containing timestamps and narrative functions, for example:

[0295]

[0296] The key frame mapping module maps the timeline segmentation scheme with the key frame events and emotional peaks in the semantic analysis results to generate a key frame mapping table, including the following steps:

[0297] 1. Event-emotion joint alignment:

[0298] Extract the event time axis (such as "product display" events at 10-15s) and emotional intensity curve (such as emotional peak at 12s with intensity 0.8) from the semantic analysis results, and align them with the timeline segmentation through the dynamic time warping (DTW) algorithm. For example, force the high-emotional-peak event to be mapped to the golden position (middle of the segment) in the main body segment.

[0299] 2. Visual saliency fusion:

[0300] Combine visual saliency maps in the annotated material (e.g., face of the person, product close-up area), and mark "visual focus windows" within the time axis segments. For example, set a focus window every 2s in the main segment, and prioritize mapping keyframes containing product close-ups.

[0301] 3. Mapping table generation strategy:

[0302] Adopt a three-layer mapping mechanism:

[0303] Basic layer: map by event type (e.g., "discount information" is mapped to the end segment);

[0304] Optimization layer: adjust mapping position according to emotional intensity (emotion peak events are offset to the golden points of the segments);

[0305] Constraint layer: avoid keyframe overlap (e.g., two high saliency events are separated by ≥1s).

[0306] Output a timestamped keyframe mapping table, for example:

[0307]

[0308] The shot timing arrangement module performs shot selection and timing logic arrangement based on the keyframe mapping table and the annotated material, and generates a shot arrangement sequence, including the following steps:

[0309] 1. Shot candidate set generation:

[0310] Extract candidate shots from the annotated material according to the time window of the keyframe mapping table:

[0311] Content matching: search for shots containing mapped events (e.g., "lipstick color testing" corresponds to close-up shots);

[0312] Quality filtering: eliminate shots with many noise points and blurring (filtered by VMAF≥80);

[0313] Diversity control: ensure that the candidate set contains shots of different perspectives (close-up / medium shot / panorama) to avoid visual monotony.

[0314] 2. Timing logic modeling:

[0315] Build a shot timing graph model, with nodes as candidate shots and edge weights determined by three factors:

[0316] Semantic coherence: calculate the semantic similarity between shots (e.g., the correlation between "product demonstration" and "use effect" shots) using BERT;

[0317] Visual fluency: Calculate the motion vector difference between adjacent shots using optical flow to avoid drastic angle changes.

[0318] Emotional consistency: Ensure smooth changes in shot emotional intensity (e.g., gradually increase from 0.5 to 0.8 instead of abruptly changing).

[0319] 3. Sequence optimization algorithm:

[0320] Use a sequence optimization algorithm based on simulated annealing, with the objective function:

[0321] Maximize (semantic coherence × 0.4 + visual fluency × 0.3 + emotional consistency × 0.3)

[0322] When processing long videos (> 60s), introduce a hierarchical scheduling strategy: first segment and schedule by minutes, then fine-tune the global timing, improving computational efficiency.

[0323] 4. Sequence output:

[0324] Generate a sequence output containing shot ID, entry / exit time, and transition type, for example:

[0325]

[0326] The editing decision generation module performs shot sequence point marking, transition type definition, and parameter configuration to generate editing decision instructions, including the following steps:

[0327] 1. Intelligent cut point marking:

[0328] Detect cut points based on shot content changes:

[0329] Content mutation: Mark shot switching points (pixel difference > 30%) using inter-frame difference method;

[0330] Semantic boundary: Force mark cut points at the end of events (e.g., "product demonstration" event tail frame);

[0331] Rhythm adaptation: In fast-paced editing, control the cut point interval to 0.5-1s, and in slow-paced editing, relax to 2-3s.

[0332] 2. Dynamic transition type definition:

[0333] Build a transition strategy knowledge base and select transition types based on three conditions:

[0334] Shot relationship: Use "fade in and out" for similar scenes and "wipe" for scenes with large differences;

[0335] Emotional intensity: Use "fast cut" for high emotional segments and "dissolve" for low emotional segments;

[0336] User preference: remember user's common transition (e.g. 80% of a user uses "fast cut") according to historical data.

[0337] Transition parameter (e.g. duration) is dynamically adjusted according to shot length: transition duration is 1s for long shot (>5s), and is shortened to 0.3s for short shot (<2s).

[0338] 3. Decision instruction generation:

[0339] Integrate the editing points, transition information and parameter configuration to generate executable editing decision instructions, for example:

[0340]

[0341] As can be seen from the above, the narrative structure output by the timeline segmentation module drives the definition of the time window of the key frame mapping, the mapped key frame guides the candidate set screening of the shot timing arrangement, and finally the editing decision generation module converts the arrangement sequence into executable instructions. If the shot semantics are not coherent in the arrangement (e.g. "product demonstration" followed by unrelated scenes), the system will backtrack to the key frame mapping module to readjust the event time window, forming a closed-loop optimization of "segmentation → mapping → arrangement → decision". This process makes the editing logic conform to the narrative rules, and compared with manual arrangement, the efficiency is improved by 15 times, and in the scenes of e-commerce promotion, brand promotion, etc., the user click rate is improved by 22%.

[0342] Optionally, an automatic editing execution unit is configured to perform automatic editing processing on the standardized editing materials according to the editing decision instructions to generate a preliminary mixed video, specifically including the following steps:

[0343] A key frame intelligent extraction subunit is configured to perform visual saliency and emotional peak detection processing on the semantic annotation materials to generate a key frame set;

[0344] A shot segmentation and reorganization subunit is configured to perform shot segmentation and timing reorganization processing on the key frame set and the editing decision instructions to generate a reorganized shot sequence;

[0345] A transition effect automatic generation subunit is configured to perform intelligent matching and generation processing on the reorganized shot sequence to generate a shot sequence with transition effects;

[0346] A multi-track audio fusion subunit is configured to perform multi-track fusion processing on the shot sequence with transition effects to generate a preliminary mixed video.

[0347] Optionally, the key frame intelligent extraction subunit includes:

[0348] A visual saliency calculation module is configured to perform pixel-level visual saliency region detection processing on the semantic annotation materials to generate a saliency map sequence;

[0349] a sentiment feature fusion module configured to perform sentiment feature weighted fusion processing on the saliency map sequence and the semantic parsing result to generate a sentiment saliency map;

[0350] a key frame candidate generation module configured to perform spatiotemporal continuity analysis and key frame candidate extraction processing on the sentiment saliency map to generate a key frame candidate set;

[0351] a key frame screening module configured to perform semantic representativeness and visual diversity evaluation processing on the key frame candidate set to generate a key frame set.

[0352] The visual saliency calculation module performs pixel-level visual saliency region detection on the semantic annotation material, and generates a saliency map sequence, including the following steps:

[0353] 1. Multi-scale saliency detection network:

[0354] An improved BASNet (boundary-aware saliency network) is used as the basic architecture, and a nested U-Net structure is used to realize pixel-level saliency prediction. The network input is a 3-channel RGB frame (resolution 1280x720), and the output is a saliency probability map (0-1 value, the higher the value, the more significant) with the same size as the input.

[0355] A cross-layer attention mechanism is introduced to extract high-level semantic features (such as "human face" and "product main body") while preserving low-level edge details (such as hair and texture), solving the problem of traditional saliency models that the main body is complete but the edge is blurred.

[0356] 2. Spatiotemporal consistency enhancement:

[0357] For a video frame sequence, 3D convolution (3x3x3 kernel) is used to process 5 consecutive frames to capture the saliency trajectory of moving objects (such as moving people), avoiding saliency fluctuations caused by motion blur in single-frame detection.

[0358] A dynamic threshold method is used to handle different scenes: a fixed threshold (0.5) is used to segment the saliency region in indoor scenes, and an adaptive threshold (based on the global saliency mean + standard deviation) is used to dynamically adjust in complex outdoor scenes, improving the accuracy of main body detection in complex backgrounds.

[0359] 3. Output format:

[0360] A frame-by-frame saliency map sequence is generated, with each map being a HxWx1 floating-point matrix, for example, a 1080P frame corresponds to a 1920x1080x1 map, where the area with a value >0.5 is the salient area.

[0361] The emotional feature fusion module fuses the saliency map sequence and the emotional features of the semantic analysis results, and generates an emotional saliency map, including the following steps:

[0362] 1. Emotional feature extraction:

[0363] Obtain the emotional intensity curve (0-1 value, such as the "excitement" emotional intensity over time) from the semantic analysis results, and extract the audio emotional features (such as the emotional probability distribution mapped by the MFCC features).

[0364] For video frames, use the EmoNet model to extract visual emotional features, focusing on analyzing facial expressions (such as smiling, surprised) and color emotions (such as red -> excitement, blue -> calm), and generate frame-level emotional vectors (8-dimensional emotional space).

[0365] 2. Weighted fusion strategy:

[0366] Design a three-layer fusion weight model:

[0367] Bottom weight: saliency map accounts for 60% (basic visual importance);

[0368] Middle weight: visual emotional features account for 30% (such as high emotional frames enhance saliency);

[0369] High-level weight: audio emotional features account for 10% (such as music climax corresponding to the saliency of the picture is improved).

[0370] Introduce attention mechanism, dynamically adjust the weight according to the event type in semantic analysis: in "promotion" event, the visual saliency weight is increased to 70%, and the emotional weight is reduced to 20%, to ensure that the product area is prioritized.

[0371] 3. Emotional saliency map generation:

[0372] Calculate the weighted fusion value for each frame: `Emo_Saliency = Saliency × W1 + Visual_Emotion × W2 + Audio_Emotion × W3`, generate 0-1 emotional saliency value, and finally output the emotional saliency map sequence with the same size as the original frame.

[0373] The key frame candidate generation module performs spatiotemporal analysis on the emotional saliency map, and extracts the key frame candidate set, including the following steps:

[0374] 1. Spatiotemporal continuity analysis:

[0375] Use the optical flow method to calculate the motion vector of adjacent frames, and construct the motion trajectory of the salient region. If a region maintains high saliency for more than 3 consecutive frames and the motion trajectory is smooth, it is marked as a "continuous salient region" to avoid misjudgment caused by transient saliency fluctuations.

[0376] With the time pyramid structure, we analyze the trend of saliency changes at different time scales (1s, 5s, 10s) to capture long-time salient events (e.g., continuous product display) and short-time peaks (e.g., promotional slogan flashing).

[0377] 2. Candidate frame extraction algorithm:

[0378] Adopting a "peak detection + uniform sampling" dual strategy:

[0379] Peak detection: Identify local maximum points in the emotional saliency curve (e.g., the highest value in consecutive 3 frames) to ensure that key event frames (e.g., product close-up, emotional climax) are extracted;

[0380] Uniform sampling: For flat sections without obvious peaks, sample at fixed intervals (e.g., every 10 seconds) to ensure content coverage.

[0381] De-duplicate the candidate frames: Calculate the visual similarity between candidate frames (based on cosine distance of feature vectors), and only keep the frame with the highest emotional saliency if the similarity is >0.8.

[0382] 3. Candidate set generation:

[0383] Output the candidate set containing frame index, emotional saliency value, and motion intensity, for example:

[0384]

[0385] The key frame selection module evaluates the semantic representativeness and visual diversity of the candidate set, and generates the final key frame set, including the following steps:

[0386] 1. Semantic representativeness evaluation:

[0387] Calculate the similarity between the visual features of the candidate frames (e.g., image features extracted by ResNet) and the event vectors in the semantic analysis results, and prefer to keep frames with high matching degree to key events such as "product display" and "promotional slogan" (cosine similarity >0.7).

[0388] Introduce knowledge graph constraints: If the candidate frame contains core entities in the knowledge graph (e.g., "lipstick" and "limited-time discount"), its semantic representativeness score is increased by 30%.

[0389] 2. Visual diversity evaluation:

[0390] Use clustering algorithm to cluster the visual features of candidate frames, and each cluster center represents a visual mode (e.g., "product close-up" and "panoramic scene"), ensuring that key frames cover at least 3 different clusters and avoiding visual repetition.

[0391] Calculate the inter-frame difference: use pHash to calculate the difference between the candidate frame and the selected key frame. Frames with a difference <0.5 are filtered to ensure diversity (a difference of 1 indicates complete difference).

[0392] 3. Multi-objective optimization screening:

[0393] Construct a multi-objective optimization function:

[0394] Score = semantic representativeness x 0.5 + visual diversity x 0.3 + emotional significance x 0.2

[0395] Use the non-dominated sorting genetic algorithm (NSGA-II) to solve the optimal solution, balance the three dimensions of semantics, vision, and emotion, and finally select the TOP-N key frames (N = 5%-10% of the total number of frames).

[0396] As can be seen from the above, the atlas output by the visual saliency calculation module provides a visual basis for emotional fusion, and the fused emotional saliency atlas guides the extraction of candidate frames, and the screening module further optimizes the candidate set. If the semantic coverage of the key frames after screening is insufficient (such as missing important events), the system will backtrack to the candidate generation module and adjust the sampling strategy (such as reducing the uniform sampling interval) to form a closed loop of "saliency calculation → emotional fusion → candidate generation → screening optimization". This process makes the key frame extraction accuracy reach 92.4%, which is 27% higher than traditional saliency methods, especially in e-commerce promotion videos, the recall rate of product key frames is increased from 68% to 95%.

[0397] Optionally, the shot segmentation and reorganization subunit comprises:

[0398] A shot boundary detection module for detecting shot switching points in the semantically annotated material to generate a list of shot boundary time points;

[0399] A shot semantic annotation module for annotating the semantic theme and emotional label of the shots corresponding to the list of shot boundary time points to generate a set of semantic shots;

[0400] A clip decision mapping module for mapping and matching the semantic theme and time axis of the set of semantic shots and the clip decision instructions to generate a set of mapped shots;

[0401] A timing reorganization optimization module for optimizing the narrative logic and timing of the set of mapped shots to generate a sequence of reorganized shots.

[0402] When the shot boundary detection module detects the shot switching points in the semantically annotated material to generate a list of shot boundary time points, the following steps are included:

[0403] 1. Visual feature difference detection:

[0404] Dual-threshold inter-frame difference method: Calculate the difference in luminance histogram and pixel gradient of adjacent frames. When both exceed the threshold (luminance difference > 20% and gradient difference > 30%), mark as potential boundary points.

[0405] Introduce Gaussian pyramid multi-scale analysis, fuse frame difference results at different resolution layers (such as 1 / 4, 1 / 2, original size), avoid misjudgment at single scale (such as false boundary caused by fast motion).

[0406] 2. Semantic mutation detection:

[0407] Use pre-trained CLIP model to calculate semantic similarity of adjacent frames. When the cosine distance of semantic vectors > 0.6, determine as semantic mutation boundary (such as scene switching from indoor to outdoor).

[0408] Combine entity changes in semantic annotation materials, if the subject entity (such as "lipstick" changes to "foundation") or scene label (such as "close-up" changes to "panorama") mutation is detected, force mark as shot boundary.

[0409] 3. Boundary point post-processing:

[0410] Use sliding window smoothing (window size 5 frames) to eliminate dense false boundaries and preserve real boundary points; insert artificial heuristic boundaries (such as forced segmentation every 20 seconds) for long shots (> 30 seconds) to improve subsequent editing flexibility.

[0411] Output boundary time point list, format is `[00:00:05.234,00:00:12.567,...]`, accurate to millisecond level.

[0412] Shot semantic annotation module annotates the semantic theme and emotional label of the shot corresponding to the shot boundary, generates semantic shot set, including the following steps:

[0413] 1. Multi-modal feature extraction:

[0414] Visual features: use Swin Transformer to extract global semantic features of key frames (such as "product display" "promotion scene"), combine YOLOv8 to detect entity categories (such as "lipstick" "discount label");

[0415] Audio features: extract MFCC features and analyze tone changes through BiLSTM to identify emotional tendencies such as "excited" "urgent";

[0416] Text features: use BERT to encode semantic vectors for OCR text in the shot (such as promotional slogans).

[0417] 2. Semantic theme classification:

[0418] Three-tier classifier is constructed: the first tier identifies the basic scene (indoor / outdoor / close-up), the second tier identifies the event type (product display / promotion), and the third tier identifies the specific entity (such as "lipstick test color" "limited-time discount"). The classifier is fine-tuned based on 100,000+ labeled shots, with an F1 value of 0.92.

[0419] Knowledge graph enhancement is introduced: the identified entity (such as "lipstick") is associated with the "cosmetic product" node in the graph, enriching the semantic hierarchy (such as "lipstick → cosmetics → fast-moving consumer goods").

[0420] 3. Emotional label generation:

[0421] Fusion of visual emotion (facial expression, color) and audio emotion (tone, rhythm), weighted by attention mechanism (60% visual, 40% audio), output "excited" "warm" and other 8 types of emotional labels and intensity values (0-1).

[0422] Output semantic shot structure:

[0423]

[0424]

[0425] The editing decision mapping module maps the semantic shots to the editing decision instructions to generate a set of mapped shots, including the following steps:

[0426] 1. Decision instruction analysis:

[0427] Extract time axis segmentation (opening / subject / ending), key events (such as "product close-up" "promotion information") and emotional requirements (such as "fast rhythm" "high excitement level") from the editing decision instructions, and construct the target semantic vector (such as "subject segment + product display + excitement").

[0428] 2. Semantic matching algorithm:

[0429] Three-tier matching strategy is adopted:

[0430] Basic matching: calculate the keyword overlap rate of shot semantics and decision instruction (such as "product display" matching degree);

[0431] Emotional matching: calculate the difference between shot emotional intensity and decision requirement (such as requiring excitement level > 0.6, and selecting corresponding shots);

[0432] Structural matching: according to the time axis segmentation constraint (such as the opening segment only matches "brand logo" type shots).

[0433] Encode the shot semantics and decision instructions into a unified vector space using a Transformer encoder. Sort the vectors by cosine similarity (threshold > 0.6).

[0434] 3. Mapping optimization mechanism:

[0435] When handling conflict situations (e.g., high semantic match but poor emotional alignment), enable reinforcement learning for reordering: prioritize shots with emotional alignment (weight 0.7), followed by semantic considerations (0.3).

[0436] Output mapping results:

[0437]

[0438] The temporal reorganization optimization module optimizes the narrative logic and temporal reorganization of the mapped shots, generating a reorganized shot sequence. The process includes the following steps:

[0439] 1. Narrative logic modeling:

[0440] Construct an advertisement narrative graph and define narrative patterns such as "problem-solution" and "feature-advantage-benefit." Each pattern corresponds to temporal constraints for shot types (e.g., a "product problem" shot must be followed by a "solution" shot).

[0441] Use a state machine to represent the narrative flow, with states as semantic categories and transition conditions as semantic relevance (e.g., a "product demonstration" can transition to "use effect" or "discount information").

[0442] 2. Temporal optimization algorithm:

[0443] Construct a shot temporal graph based on Graph Neural Networks (GNN), with nodes representing shots and edge weights representing semantic coherence (computed by BERT), visual fluency (computed by optical flow to calculate motion differences), and emotional consistency (absolute difference).

[0444] Use a simulated annealing algorithm to solve the optimal sequence, with the objective function:

[0445] Maximize (narrative logic compliance x 0.4 + visual fluency x 0.3 + emotional consistency x 0.3)

[0446] When processing long videos, first optimize locally by narrative paragraphs (e.g., every 10 seconds), then perform global fine-tuning to improve computational efficiency.

[0447] 3. Reorganized sequence generation:

[0448] Output the reorganized sequence containing shot order and transition methods:

[0449]

[0450] In summary, the boundary points output by the lens boundary detection module define the lens range, the semantic labeling module assigns semantics and emotional attributes to each lens, the mapping module aligns the lens with the decision instructions, and the reorganization optimization module generates the final sequence. If the reorganization finds a narrative logic break (such as "discount information" followed by an unrelated lens), the system will backtrack to the mapping module to re-screen the lens, forming a closed loop of "boundary detection → semantic labeling → mapping matching → reorganization optimization". This process improves the narrative rationality of lens reorganization by 35%, and the user click rate is 28% higher than random arrangement. In e-commerce promotion videos, the optimization of the display order of product conversion-related lenses improves the conversion rate by 19%.

[0451] Optionally, the transition effect automatic generation subunit comprises:

[0452] A transition type prediction module is configured to perform semantic similarity and visual difference analysis on adjacent lenses of the reorganized lens sequence to generate a transition type prediction result.

[0453] An effect parameter generation module is configured to generate effect parameters (duration / direction / effect style) based on the transition type prediction result and the editing decision instruction to generate a set of effect parameters.

[0454] A transition effect generation module is configured to instantiate and render the set of effect parameters to generate transition effect materials.

[0455] A transition fluency optimization module is configured to optimize the timing fluency and visual continuity of the transition effect materials and the reorganized lens sequence to generate a lens sequence with transitions.

[0456] When the transition type prediction module performs semantic similarity and visual difference analysis on adjacent lenses of the reorganized lens sequence to generate a transition type prediction result, the following steps are included:

[0457] 1. Multi-modal feature extraction:

[0458] Visual features: Use the CLIP model to extract the visual semantic features of the key frames of adjacent lenses, calculate the visual similarity (such as cosine distance) by comparing feature vectors; at the same time, use the optical flow method to calculate the inter-frame motion vector difference, quantify the visual mutation degree of lens switching.

[0459] Semantic features: Extract topic labels (such as "product display" and "discount information") and emotional intensity from lens semantic labeling, calculate semantic coherence (such as whether the topics of adjacent lenses are related) by BERT.

[0460] 2. Transition type classification model:

[0461] Three-layer classifier is constructed: the first layer judges the transition type category (shear / dissolve / slide / special effect), the second layer subdivides the specific type (such as "fast cut" "fade in and out" "left and right slide"), and the third layer optimizes the parameter tendency (such as the direction preference of special effect transition). The model is trained based on 500,000+ labeled shots, and the F1 value reaches 0.91.

[0462] Attention mechanism is introduced to dynamically adjust feature weights according to shot type: increase visual difference weight (60%) between product close-up shots, and increase semantic weight (70%) between scene switching shots.

[0463] 3. Prediction result generation:

[0464] Output transition type and confidence, for example: `{"type":"cut","confidence":0.85}` (fast cut), `{"type":"fade","confidence":0.72}` (fade in and out), support "automatic" mode decided by the model or "manual" mode according to user preference.

[0465] When the special effect parameter generation module generates a set of special effect parameters (duration / direction / style) based on the transition type prediction result and the editing decision instruction, the following steps are included:

[0466] 1. Parameter rule engine:

[0467] Build a transition parameter knowledge base containing 200+ preset rules:

[0468] Fast cut (cut): fixed duration 0.1s, no direction, default style is none;

[0469] Fade in and out (fade): duration dynamically adjusted according to emotional intensity (high emotion → 0.5s, low emotion → 1s), direction default "none";

[0470] Slide (slide): direction consistent with shot motion direction (detected by optical flow method), duration positively related to shot duration (long shot → 1.2s).

[0471] 2. User intent adaptation:

[0472] Parse parameter constraints in editing decision instructions (such as "fast pace" "technology feature effect"), and adjust parameters through reinforcement learning:

[0473] Under the demand of "fast pace", all transition durations are shortened by 20%;

[0474] "Technology" demand prefers "matrix switching" "particle effect" and other styles, and the direction is set to "diagonal line".

[0475] 3. Parameter conflict resolution:

[0476] When dealing with conflicting parameters (e.g., "slow pace" vs. "fast cut"), prioritize as follows: user instructions > emotional intensity > visual features. For example, if the user insists on a "fast cut," ignore the pace matching degree and only adjust the duration to 0.3s as a compromise value.

[0477] Output parameter set: `{"duration":0.5s,"direction":"left_to_right","style":"glow_transition"}`.

[0478] The transition effect generation module instantiates and renders the effect parameter set, and when generating the transition effect material, the following steps are included:

[0479] 1. Effect template library call:

[0480] Maintain 100+ transition effect templates, classified by type (e.g., basic transitions, e-commerce effects, technology effects), each template containing pre-rendered resources and configurable parameter interfaces. For example, the "shining gold" effect template supports adjusting light intensity, particle density, and other parameters.

[0481] 2. Real-time rendering engine:

[0482] Based on a GPU-accelerated real-time rendering framework (e.g., a lightweight version of Unity / Unreal Engine), dynamically generate effects after receiving parameters:

[0483] For "fade-in and fade-out" transitions, calculate the pixel blending weight curve of adjacent shots;

[0484] For "slide" transitions, generate layer movement animations with motion blur;

[0485] For special effect transitions, instantiate 3D models (e.g., rotating cubes) and bind parameters (rotation speed, material color).

[0486] 3. Material generation strategy:

[0487] Adopt a "pre-rendering + real-time synthesis" mode: commonly used transitions are pre-rendered as transparent channel video segments (e.g., "fast cut" and "fade-in and fade-out"), and special effects are generated in real time. The generated transition material resolution is consistent with the original video, and the frame rate matches the project settings (e.g., 30fps).

[0488] The transition fluency optimization module optimizes the timing fluency and visual continuity of transition effect materials and recombined shot sequences, including the following steps:

[0489] 1. Timing smoothing processing:

[0490] Optical flow frame interpolation: For dynamic transitions such as "slide" and "zoom", calculate the motion vector between adjacent shots to generate intermediate transition frames (e.g. insert 2 frames every 0.1s) to eliminate the feeling of stuttering.

[0491] Audio-video synchronization: Analyze the audio energy changes before and after the transition, adjust the transition duration to synchronize the audio transition with the visual transition (e.g. complete the transition at the drum point of the music).

[0492] 2. Visual consistency optimization:

[0493] Color matching: Use 3D LUT to unify the color of the shots before and after the transition, avoid color mutation (e.g. add a gradual color filter when suddenly changing from cold to warm color).

[0494] Brightness equalization: Through histogram matching, ensure that the brightness distribution of the frames before and after the transition is consistent, eliminate visual flicker (e.g. add brightness transition effects when transitioning from dark to bright scenes).

[0495] 3. Smoothness evaluation and correction:

[0496] Use VMAF-FR (full reference) index to evaluate the transition quality, if the score is <85, automatically adjust the parameters:

[0497] Stutter → increase the number of frame interpolation;

[0498] Color jump → strengthen LUT matching;

[0499] Audio and video out of sync → recalculate the audio synchronization point.

[0500] As can be seen from the above, the type guidance parameter generated by the transition type prediction module, the parameter drives the special effect rendering, and the optimization module ensures the smoothness. If there is still visual inconsistency (such as large color difference) after optimization, the system will backtrack to the parameter generation module to adjust the LUT parameters, forming a "prediction → parameter → generation → optimization" closed loop. This process improves the efficiency of transition special effect generation by 40%, and the user's subjective satisfaction score is improved by 25% compared with manual production. Especially in e-commerce promotion videos, the smoothness of fast-paced transitions is significantly improved, and the completion rate of watching is increased by 18%.

[0501] Optionally, the multi-track audio fusion sub-unit comprises:

[0502] An audio feature extraction module for multi-track feature extraction processing of voice, sound effects and background music of the shot sequence with transitions to generate an audio feature set;

[0503] An audio scene matching module for audio scene (promotion / warm / dynamic) matching processing according to the audio feature set and the semantic analysis result to generate a candidate audio material set;

[0504] An audio parameter adjustment module is configured to perform volume, rhythm, and semantic synchronization adjustment processing on the candidate audio material set to generate adjusted audio materials.

[0505] A multi-track mixing and rendering module is configured to perform multi-track audio mixing and audio-visual synchronization rendering processing on the adjusted audio materials and the shot sequence with transitions to generate a preliminary mixed and cut video.

[0506] The audio feature extraction module performs multi-track feature extraction of speech, sound effects, and background music on the shot sequence with transitions, and generates an audio feature set, including the following steps:

[0507] 1. Multi-track separation technology:

[0508] Wave-U-Net network is used to separate the original audio, and the input mixed audio (such as vocals + background music + environmental sound) is decomposed into independent speech track, background music track, and sound effect track. The network learns the spectral features of different audio sources through training, and the separation accuracy reaches 90% (SDR≥15dB).

[0509] Time-domain and frequency-domain feature extraction is performed on each track after separation:

[0510] Speech track: extract pitch, speech rate, energy envelope, and mel-frequency cepstral coefficient (MFCC);

[0511] Background music track: extract rhythm (BPM), tonality, and emotion label (such as "happy" and "relaxed");

[0512] Sound effect track: extract sound pressure level (SPL), spectral distribution, and duration.

[0513] 2. Audio scene analysis:

[0514] Combine shot semantic annotation (such as "product demonstration" and "discount information") to classify the scene of each audio. For example, when the shot semantic is detected as "limited-time discount", focus on analyzing the keywords (such as "buy immediately") in the speech track and their audio energy changes.

[0515] Build an audio emotion classifier based on the ResNet architecture to identify the emotion of the audio segment (such as "excited" and "calm"), with a classification accuracy of 85%.

[0516] 3. Feature set generation:

[0517] Output a structured feature set, for example:

[0518]

[0519]

[0520] The audio scene matching module performs audio scene matching based on the audio feature set and the semantic analysis result, and generates a candidate audio material set, including the following steps:

[0521] 1. Scene template library:

[0522] Construct a knowledge base containing 100+ audio scene templates, each associated with a specific combination of audio features:

[0523] Promotion scene: high-energy background music (BPM ≥ 120), clear speech (SNR ≥ 15 dB), high-frequency sound effects (such as countdown prompt sound);

[0524] Warm scene: soothing background music (BPM ≤ 80), soft voice, low-frequency environmental sound (such as soft wind);

[0525] Dynamic scene: strong rhythm drum (BPM ≥ 130), dynamic voice change, high-frequency percussion sound effect.

[0526] 2. Similarity matching algorithm:

[0527] Calculate the similarity between the current audio features and the scene template, using weighted distance measurement:

[0528] Voice feature weight 0.4 (keyword matching degree + pitch stability);

[0529] Background music weight 0.4 (BPM matching degree + emotional consistency);

[0530] Sound effect weight 0.2 (type matching degree + time synchronization).

[0531] 3. Candidate material screening:

[0532] Retrieve the highest matching candidate audio from the audio material library:

[0533] For promotion scenes, prefer music containing "promotion horn" sound effects and lively rhythms;

[0534] For warm scenes, filter background music dominated by piano or string instruments;

[0535] For dynamic scenes, match electronic dance music or rock-style music.

[0536] Output candidate set:

[0537]

[0538] The audio parameter adjustment module adjusts the volume, rhythm, and semantic synchronization of the candidate audio material, and generates the adjusted audio material, including the following steps:

[0539] 1. Volume balance control:

[0540] Dynamic adjustment of individual track volumes based on voice activity detection (VAD):

[0541] Automatic 3-5dB reduction of background music when valid speech is present in the speech track;

[0542] Temporary 2dB boost of speech track volume when keywords (e.g. "Attention" "Look here") are detected;

[0543] Dynamic adjustment of gain for sound effect tracks based on importance ranking (e.g. "Promotional horn" has higher priority than ordinary click sound).

[0544] 2. Rhythm synchronization optimization:

[0545] Analysis of video transition timing (e.g. shot cut points) to align drum hits or melody climaxes of background music with transitions. For example, when a fast cut transition is detected, ensure that the music has a rhythmic accent at the moment of the transition.

[0546] Time stretching of speech tracks to adjust speech rate without changing pitch, matching video rhythm (e.g. fast-paced video corresponds to speech rate ≥ 160 words / minute).

[0547] 3. Semantic synchronization enhancement:

[0548] Adjusting audio parameters according to shot semantics:

[0549] In "product close-up" shots, enhance environmental sound effects (e.g. product packaging sound);

[0550] In "discount information" shots, add low-frequency emphasis sound (e.g. bass drum);

[0551] In emotional climax shots, increase the dynamic range of background music (e.g. from weak to strong crescendo).

[0552] The multi-track mixing and rendering module performs multi-track mixing and audio-visual synchronization rendering on the adjusted audio materials and transitioned shot sequences to generate a preliminary mixed video, including the following steps:

[0553] 1. Multi-track mixing engine:

[0554] Using 3D audio spatialization technology to assign virtual spatial positions for different audio tracks:

[0555] Speech track placed in front (0° azimuth), 0.5m away;

[0556] Background music track distributed around (360°), 2m away;

[0557] Special effect sound is positioned according to semantics (e.g. sound effect of "product on the left" is placed in the left channel).

[0558] 2. Audio-visual synchronization technology:

[0559] Calculate the motion speed between video frames based on the optical flow method, and adjust the audio playback speed synchronously. For example, when the video is played at an accelerated speed, the audio is time-compressed through the WSOLA algorithm to maintain the pitch unchanged.

[0560] For transition effects such as fade-in and fade-out, apply audio fade-in and fade-out effects synchronously to ensure consistent visual and auditory transitions.

[0561] 3. Real-time rendering and quality control:

[0562] Use the AES67 standard for multi-track audio transmission to ensure low-latency (<1ms) mixing;

[0563] Perform dynamic range compression (DRC) on the output audio to limit peak values (-3dBFS) and boost average volume (LUFS = -16);

[0564] Generate a preview version and evaluate speech clarity through the Short-Time Objective Intelligibility Index (STOI). If STOI < 0.9, automatically adjust the speech enhancement parameters.

[0565] In summary, the audio feature extraction module provides data basis for scene matching, the matching result guides parameter adjustment, and the adjusted audio is mixed and rendered synchronously with the video. If the rendered audio is found to be out of sync with the video (e.g., speech and lip movements are misaligned), the system will backtrack to the parameter adjustment module to correct the time offset, forming a closed loop of "extraction → matching → adjustment → rendering". This process improves audio production efficiency by 50%, and the audio-visual synchronization accuracy rate reaches 98%. In e-commerce videos, the improvement in speech clarity increases audience comprehension by 22%, and product conversion rate by 15%.

[0566] Optionally, an effect optimization output unit is configured to perform visual effect enhancement and format optimization on the preliminary mixed and edited video to generate a target mixed and edited video, including the following steps:

[0567] A visual effect enhancement subunit is configured to perform color correction, sharpening, and special effect superposition processing on the preliminary mixed and edited video to generate a visual enhancement video;

[0568] A subtitle intelligent generation and positioning subunit is configured to perform speech recognition and subtitle generation on the visual enhancement video, and perform subtitle positioning processing in combination with picture area detection to generate a video with subtitles;

[0569] A multi-terminal adaptation optimization subunit is configured to perform multi-terminal adaptation optimization processing on the video with subtitles in terms of resolution, code rate, and format to generate an adaptation optimized video;

[0570] The quality evaluation and output subunit is configured to evaluate the fluency, picture quality, and semantic consistency of the adaptive optimization video, and perform optimization processing according to the evaluation result to generate a target mixed video.

[0571] Optionally, the visual effect enhancement subunit comprises:

[0572] The color space conversion module is configured to perform color space (RGB / YCbCr) conversion and color gamut normalization processing on the preliminary mixed video to generate a color normalized video.

[0573] The color correction module is configured to perform brightness / contrast adjustment, color balance, and stylized filter processing on the color normalized video to generate a color corrected video.

[0574] The super-resolution module is configured to perform resolution enhancement and detail recovery processing on the color corrected video to generate a super-resolution video.

[0575] The visual effect superimposition module is configured to perform dynamic sticker, special effect element, and subtitle undercoat superimposition processing on the super-resolution video to generate a visual enhancement video.

[0576] When the color space conversion module performs color space (RGB / YCbCr) conversion and color gamut normalization on the preliminary mixed video to generate a color normalized video, the following steps are included:

[0577] 1. Color space conversion engine:

[0578] Matrix transformation is used to realize bidirectional conversion between RGB and YCbCr spaces, for example, RGB is converted to YCbCr through the following formula:

[0579] Y=0.299R+0.587G+0.114B

[0580] Cb=-0.1687R0.3313G+0.5B+128

[0581] Cr=0.5R0.4187G0.0813B+128

[0582] Conversion parameters of different standards such as BT.601 and BT.709 are supported, and the conversion parameters are automatically selected according to the target platform of the video (such as BT.709 for television and sRGB for mobile devices).

[0583] 2. Color gamut normalization processing:

[0584] The color gamut space of the input video is detected (such as Adobe RGB, DCI-P3), and the color gamut mapping algorithm is used to compress it to the target color gamut (such as sRGB):

[0585] For colors beyond the target gamut, use saturation compression (keep hue unchanged, reduce saturation);

[0586] For HDR videos, use PQ curve or HLG curve for SDR mapping, preserving brightness information while avoiding overexposure.

[0587] 3. Output control:

[0588] When generating color-standardized videos, preserve the color semantic information of the original video, such as ensuring consistent product color (e.g., red lipstick) on different devices in e-commerce videos, with a color difference ΔE ≤ 3.

[0589] The color correction module performs brightness / contrast adjustment, color balance, and stylized filter processing on the color-standardized video to generate a color-corrected video, including the following steps:

[0590] 1. Adaptive brightness and contrast adjustment:

[0591] Use local histogram equalization to divide the video frame into 8x8 pixel blocks, calculate the histogram for each block and stretch it to the target range, and enhance dark details (e.g., texture in product shadows).

[0592] Adjust the brightness curve dynamically through Gamma correction, such as increasing the Gamma value of the product area in e-commerce videos to 1.8 to make the product brighter.

[0593] 2. 3D color balance processing:

[0594] Construct a 3D LUT (Lookup Table) for color mapping, supporting the following adjustments:

[0595] Skin color optimization: Detect the face area and map the YCrCb values in the skin color range to a more natural interval.

[0596] Product color enhancement: Enhance the red channel saturation by 20% for lipstick areas in makeup videos.

[0597] Pre-set 10+ industry style LUTs (e.g., "Promotion Warm Tone" and "Technology Cool Tone"), which can be automatically applied through semantic analysis.

[0598] 3. Stylized filter generation:

[0599] Generate custom filters based on GAN networks, such as inputting "movie feel" style keywords, and the model automatically generates filters with dark corners and film grain.

[0600] Support real-time preview and parameter fine-tuning, such as adjusting filter intensity (0-100%) and application range (full frame / local).

[0601] The super-resolution module performs resolution enhancement and detail restoration on the color-corrected video. When generating a super-resolution video, the following steps are included:

[0602] 1. Multi-frame super-resolution model:

[0603] Using an improved Real-ESRGAN model, inputting 5 consecutive frames of video, aligning inter-frame motion through optical flow, and using time dimension information to improve resolution:

[0604] For static frames, use single-frame super-resolution (e.g., increase 1080P to 4K);

[0605] For dynamic frames, use multi-frame fusion to reduce motion blur (e.g., edge jaggies when a person moves).

[0606] 2. Detail restoration algorithm:

[0607] Introducing an edge-aware loss function to preserve edge details while improving resolution:

[0608] Apply a sharpening filter to product outlines (e.g., the edge of a lipstick tube) to enhance edge contrast;

[0609] Use wavelet transform to decompose high-frequency details in texture-rich areas (e.g., cloth), and then re-fuse them.

[0610] 3. Real-time processing optimization:

[0611] Use a progressive super-resolution strategy: first perform 2x super-resolution, then perform 4x super-resolution on key areas (e.g., product close-ups), balance image quality and performance, and achieve a processing speed of 20fps (4K output).

[0612] The visual effect superposition module performs superposition processing on the super-resolution video, including dynamic stickers, special effects elements, and subtitle underlines, to generate a visual enhancement video, including the following steps:

[0613] 1. Intelligent sticker positioning system:

[0614] Combine YOLOv8 to detect target positions in the video (e.g., lipstick, face), and automatically attach dynamic stickers to the target area:

[0615] For "limited-time discount" stickers, detect blank areas in the picture (e.g., the upper right corner) and place them first;

[0616] For product stickers, stick closely to the product edge (e.g., 20 pixels above the lipstick).

[0617] 2. Dynamic effect generation engine:

[0618] Generate animation effects based on the timeline, supporting the following types:

[0619] Entry animation: sticker flies in from off-screen, lasts 0.5s;

[0620] Highlight animation: promo label flashes (frequency 2Hz), lasts 3s;

[0621] Alpha channel blending technology is used to ensure that the special effects are naturally integrated with the original video, with an edge feathering radius of 2 pixels.

[0622] 3. Subtitle background optimization:

[0623] Detect the text area of the picture (through OCR) to avoid blocking the original text with the subtitles;

[0624] Automatically generate a translucent background (such as a black translucent rectangle) to improve the readability of the subtitles, and the background transparency is automatically adjusted according to the background brightness (30% transparency for light background, 70% transparency for dark background).

[0625] As can be seen from the above, the color space conversion module ensures color consistency, providing standard input for subsequent correction; the color correction module enhances visual expressiveness, the super-resolution module improves image quality details, and finally the special effect superposition module adds business required elements. If color deviation is found after superimposing special effects (such as too large color difference between stickers and background), the system will backtrack to the color correction module for re-adjustment, forming a closed loop of "conversion → correction → super-resolution → superposition". This process improves the video visual quality by 40%, and in the e-commerce scene, the product color restoration accuracy rate reaches 98%, and the special effect superposition efficiency is improved by 10 times compared with manual processing.

[0626] Optionally, the subtitle intelligent generation and positioning subunit includes:

[0627] A speech recognition module for speech-to-text and punctuation addition processing on the voice track of the visually enhanced video to generate speech text content;

[0628] A subtitle style generation module for generating style of font, color and animation effect of subtitles based on speech text content and video semantics to generate a subtitle style scheme;

[0629] A picture area analysis module for detecting text area, saliency area and safety area of the visually enhanced video to generate a picture area map;

[0630] A subtitle positioning optimization module for subtitle position planning and conflict avoidance processing on speech text content, subtitle style scheme and picture area map to generate a video with subtitles.

[0631] When the speech recognition module performs speech-to-text and punctuation addition on the voice track of the visually enhanced video to generate speech text content, the following steps are included:

[0632] 1. Multi-modal speech processing framework:

[0633] Using the Whisper-large-v3 model as the base speech recognition engine, supporting both Chinese and English, as well as 10+ dialects (e.g. Cantonese, Sichuanese), through pre-training to learn the mapping relationship between audio features and text, the word error rate (WER) is ≤5%.

[0634] Pre-processed audio noise reduction: using Wave-U-Net to remove environmental noise (such as background noise, current sound), improving the purity of speech, and the SNR after noise reduction is ≥20dB.

[0635] 2. Dynamic punctuation addition algorithm:

[0636] Based on language models (such as GPT-2) to analyze the semantic structure of speech text, automatically inserting commas, periods, and other punctuation in long sentences:

[0637] Insert commas when detecting emotional pauses (audio energy drops >3dB);

[0638] Add periods after recognizing complete semantic units (such as "buy immediately").

[0639] Support custom punctuation styles (such as using more exclamation marks for e-commerce promotions, and using strict punctuation for science popularization videos).

[0640] 3. Real-time error correction mechanism:

[0641] Build industry term libraries (such as "lipstick" and "promotion"), and post-process the recognition results:

[0642] If "lipstick color" is misrecognized as "lipstick color good", automatically correct it through semantic similarity (≥0.8);

[0643] Output timestamped text content, for example:

[0644]

[0645] The subtitle style generation module generates a subtitle style scheme based on the speech text content and video semantics, including font, color, and animation effects, including the following steps:

[0646] 1. Semantic-style mapping engine:

[0647] Build a style knowledge base and define 20+ scene style rules:

[0648] Promotion scene: bold white text + yellow outline (#FFFFFF+#FFD700), font selection "Fangzhong Huhei Jian", animation "zoom in";

[0649] Warm scene: handwritten white text (#FFFFFF), 70% transparency, animation "fade in and out";

[0650] Resolve text sentiment by BERT (e.g. "limited time offer" → excited), match corresponding style template.

[0651] 2. GAN style generation technology:

[0652] Train GAN to generate custom styles, input keywords (e.g. "tech feel") to generate blue subtitles with glowing effect (#00BFFF + external glow);

[0653] Support real-time parameter adjustment: font size (36-72px), stroke width (2-5px), animation duration (0.3-1s).

[0654] 3. Dynamic style adaptation:

[0655] Adjust style according to text length: reduce font size for long text (e.g. from 48px to 36px if more than 3 lines) to maintain readability;

[0656] Output style scheme:

[0657]

[0658] Picture area analysis module detects text area, saliency area and safety area of video, generates picture area map, including the following steps:

[0659] 1. Multi-region detection network:

[0660] Text area: use DBnet+CRNN combined model to detect printed and handwritten text in the picture, with pixel-level positioning accuracy, supporting Chinese, English and mixed text;

[0661] Saliency area: use BASNet+attention mechanism to identify visual focus (e.g. face, product subject), output saliency probability map (0-1 value);

[0662] Safety area: define 20% of screen edge as safety area (avoid subtitle blocking), center 60% as key area (prefer to keep content).

[0663] 2. Area conflict detection:

[0664] Construct area conflict matrix, calculate overlap rate of text area and saliency area:

[0665] If the overlap between subtitle candidate position and product close-up area is >30%, mark it as conflict;

[0666] Positions outside the safe area are automatically excluded (e.g. top 10% area is not placed with subtitles).

[0667] 3. Region map generation:

[0668] Output structured map, e.g.:

[0669]

[0670]

[0671] Subtitles positioning optimization module integrates speech text, style scheme and picture region map, plans subtitles position and avoids conflicts, including the following steps:

[0672] 1. Greedy positioning algorithm:

[0673] Prioritize positions within the safe area, filter in the following order:

[0674] 1. Screen bottom safe area (height 100-200px);

[0675] 2. Screen top safe area;

[0676] 3. Both sides safe area (width 100-200px);

[0677] Automatic line break for long text (≤15 words per line), line spacing set to 1.2 times the font size.

[0678] 2. Conflict avoidance strategy:

[0679] When the candidate position overlaps with the text area, move up / down 50px and retry;

[0680] If it overlaps with the saliency area >20%, adjust the subtitle transparency (to 50%) or add a semi-transparent underlay (black, transparency 70%) to improve readability.

[0681] 3. Real-time preview optimization:

[0682] Generate a preview frame with subtitles, evaluate the fusion degree of subtitles and pictures through SSIM index, if SSIM<0.8, adjust the position again;

[0683] Output the video with subtitles, the subtitle position information is synchronized to the metadata, for example:

[0684]

[0685] In summary, the text generated by the speech recognition module drives the style generation and positioning planning, and the picture area analysis provides the basis for obstacle avoidance. If the positioning finds that the subtitle blocks the key content (such as product labels), the system will backtrack to the style generation module to adjust the font size or transparency, forming a closed loop of "recognition → style → analysis → positioning". This process improves the efficiency of subtitle generation by 80%, reduces the conflict rate from 35% to 5%, and in e-commerce videos, the readability of the subtitles is improved by 27%, and the user click-through rate is increased by 19%.

[0686] Optionally, the multi-terminal adaptation optimization subunit comprises:

[0687] a terminal parameter analysis module for analyzing the resolution, bit rate and format requirements of the target output terminal (such as TikTok, YouTube, TV) to generate a terminal parameter set;

[0688] a resolution adaptation module for performing resolution scaling and black edge processing on the video with subtitles based on the terminal parameter set to generate a resolution adapted video;

[0689] a bit rate dynamic optimization module for performing bit rate and quality balancing optimization on the resolution adapted video to generate a bit rate optimized video;

[0690] a format packaging module for performing terminal format packaging and coding protocol adaptation on the bit rate optimized video to generate an adaptation optimized video.

[0691] When the terminal parameter analysis module analyzes the resolution, bit rate and format requirements of the target output terminal to generate a terminal parameter set, the following steps are included:

[0692] 1. Terminal feature library construction:

[0693] Maintain a parameter configuration library containing 100+ mainstream terminals, such as:

[0694] TikTok: Recommended resolution 720p / 1080p, bit rate 3-8Mbps, format MP4(H.264+AAC);

[0695] YouTube: Supports 4K / 8K, bit rate dynamically adjusted (4K requires 25-40Mbps), format MP4 / WEBM;

[0696] Smart TV: Resolution 1080p / 2160p, bit rate 8-15Mbps, supports Dolby Vision / HDR10.

[0697] Automatically match the parameter template through the terminal ID, such as detecting "TikTok" to load the short video optimization configuration.

[0698] 2. Dynamic parameter acquisition:

[0699] For unknown terminals (such as custom players), obtain User-Agent information through HTTP requests, and predict optimal parameters based on machine learning models.

[0700] Support user-defined parameter overrides, such as specifying the upper limit of the code rate for a specific platform.

[0701] 3. Parameter conflict handling:

[0702] When multiple parameters conflict (such as high resolution and low code rate), prioritize key indicators:

[0703] Short video platforms prioritize resolution (720p and above), and live streaming platforms prioritize frame rate (≥30fps);

[0704] Output structured parameter set:

[0705]

[0706]

[0707] When the resolution adaptation module performs resolution scaling and black border processing on videos with subtitles based on terminal parameters, the following steps are included:

[0708] 1. Intelligent scaling algorithm:

[0709] Use Bicubic algorithm for scaling, and apply Lanczos filter to edge detail areas to reduce blur;

[0710] Support non-equal scaling when the source video and target resolution ratio difference is >10%:

[0711] Prioritize preserving video main body (through saliency detection), such as human face and product area;

[0712] Intelligently crop edge areas (such as 5% on both sides) instead of directly stretching.

[0713] 2. Black border processing strategy:

[0714] When the source video ratio is inconsistent with the target, dynamically generate black borders:

[0715] Calculate black border ratio (such as adding 12.5% black borders on top and bottom when adapting 16:9 video to 4:3 terminal);

[0716] Add gradient effect (from pure black to -10% brightness) to black border area to improve visual aesthetics;

[0717] Support subtitle position recalculation to avoid subtitles being blocked by black borders.

[0718] 3. Progressive resolution adjustment:

[0719] For high resolution video (e.g. 4K), multi-stage down-sampling (4K→2K→1080p) is used to reduce information loss.

[0720] Output resolution-adaptive video, metadata contains scaling factor and black border parameters.

[0721] The code rate dynamic optimization module includes the following steps when optimizing the resolution-adaptive video for code rate-quality balance:

[0722] 1. Content-aware code rate allocation:

[0723] Divide the video frame into 8x8 macroblocks, and analyze the content complexity through CNN:

[0724] Allocate more code rate (+20%) to high motion areas (e.g. fast moving people);

[0725] Reduce code rate (-15%) for static background areas;

[0726] Increase an additional 10% code rate for text areas (e.g. subtitles, product labels) to ensure clarity.

[0727] 2. Dynamic code rate control algorithm:

[0728] Use VBR (Variable Bit Rate) encoding combined with Rate-Distortion optimization:

[0729] Increase the code rate peak value (2 times the average code rate) for scene transition frames (detected through inter-frame difference);

[0730] Reduce the code rate for smooth transition frames, maintaining the average code rate within ±10% of the target value;

[0731] Support ABR (Adaptive Bit Rate) stream generation, output multiple code rate versions (e.g. 3Mbps / 5Mbps / 8Mbps).

[0732] 3. Quality evaluation feedback:

[0733] Real-time evaluation of encoding quality through PSNR and SSIM indicators, automatically adjust parameters when SSIM < 0.95:

[0734] Reduce I-frame interval (from default 12 frames to 6 frames) to improve random access performance;

[0735] Enable code rate compensation mechanism for secondary encoding of low-quality areas.

[0736] Format packaging module includes the following steps when packaging and encoding protocol adaptation for code rate optimized video:

[0737] 1. Multi-format packaging engine:

[0738] Supports 10+ container formats including MP4, WEBM, and TS, automatically selecting the appropriate format based on the terminal requirements.

[0739] MP4 is prioritized for mobile devices (broadly compatible), while WEBM is added as an alternative for web devices;

[0740] Optimize packaging parameters:

[0741] Adjust the position of the moov atom to the beginning of the file to shorten the video loading time (<1s);

[0742] Set an appropriate segment duration (4-6 seconds) to support HTTP Live Streaming (HLS).

[0743] 2. Encoding protocol adaptation:

[0744] Video encoding supports H.264, H.265, VP9, ​​etc., and audio encoding supports AAC, OPUS, AC-3.

[0745] For iOS devices, H.264+AAC combination is forced; for Android devices, H.265 is prioritized to reduce bandwidth consumption.

[0746] For terminals that require high-quality audio (such as smart TVs), enable AC-3 5.1 channel encoding.

[0747] 3. Metadata Injection and Validation:

[0748] Automatically add key metadata:

[0749] Technical parameters such as video duration, resolution, and frame rate;

[0750] Subtitle track information (language, position), chapter markers (such as segmentation of e-commerce products);

[0751] Perform format validation before output to ensure the file conforms to terminal specifications (such as YouTube's file size limit).

[0752] In summary, terminal parameter parsing provides target configurations for subsequent modules, resolution adaptation ensures visual integrity, bitrate optimization balances bandwidth and image quality, and format encapsulation guarantees compatibility. If format incompatibility is found after encapsulation (e.g., a certain browser cannot play the video), the system will backtrack to the format encapsulation module to change the encoding protocol, forming a closed loop of "parsing → adaptation → optimization → encapsulation". This process increases the video playback success rate on various terminals to 99%, speeds up loading by 40%, and reduces bandwidth consumption by 35%. Especially in mobile weak network environments, the stuttering rate drops from 22% to 5%.

[0753] Optionally, the quality assessment and output sub-unit includes:

[0754] a picture quality objective evaluation module configured to calculate and process objective picture quality indexes such as VMAF and PSNR of the adaptation-optimized video to generate a picture quality evaluation result;

[0755] a fluency detection module configured to perform frame rate consistency and motion fluency detection processing on the adaptation-optimized video to generate a fluency evaluation result;

[0756] a semantic consistency verification module configured to perform subtitle picture semantic and audio visual semantic consistency verification processing on the adaptation-optimized video to generate a semantic evaluation result;

[0757] an optimization iteration and output module configured to perform iterative optimization processing according to the picture quality, fluency, and semantic evaluation results and output a final video.

[0758] When the picture quality objective evaluation module calculates objective picture quality indexes such as VMAF and PSNR of the adaptation-optimized video to generate a picture quality evaluation result, the following steps are included:

[0759] 1. Multi-index fusion evaluation framework:

[0760] Parallelly calculate indexes such as VMAF (video multi-method evaluation fusion), PSNR (peak signal-to-noise ratio), SSIM (structural similarity), and MS-SSIM (multi-scale structural similarity):

[0761] VMAF uses the Netflix open source model to analyze human eye perception quality through CNN, and the score range is 0-100 (≥ 90 is high quality);

[0762] PSNR measures pixel-level error, and the threshold is set to 30 dB (≥ 30 dB has no obvious distortion);

[0763] SSIM / MS-SSIM evaluates image structural similarity, and the threshold is ≥ 0.95.

[0764] Weighted fusion result: VMAF weight 0.6, SSIM weight 0.3, and PSNR weight 0.1 to form a comprehensive score.

[0765] 2. Region adaptive analysis:

[0766] Divide the video frame into a 16x16 grid, and additionally increase the weight of high semantic areas (face, text) by 20%;

[0767] Reduce the weight of transition areas (such as fade-in and fade-out) by 15% to avoid the influence of transition effects on the overall score.

[0768] 3. Defect detection engine:

[0769] Identify encoding defects such as blocking artifacts, blurring, ringing artifacts, etc.:

[0770] Blocking artifacts are detected by edge gradient mutation. If more than 5% of the area has blocking artifacts, it is determined to be a defect.

[0771] Blurring detection uses the Laplacian operator to calculate image sharpness. If it is lower than the threshold, it is marked as blurred.

[0772] Output detailed evaluation report:

[0773]

[0774] The smoothness detection module detects the consistency of video frame rate and motion smoothness. When generating the smoothness evaluation result, the following steps are included:

[0775] 1. Frame rate fluctuation analysis:

[0776] Actual frame rate is analyzed through video metadata, and frame rate fluctuation coefficient (FVF) is calculated:

[0777] Collect 100 frames of samples, calculate the standard deviation of each frame interval time, and the threshold is ≤0.01 seconds.

[0778] Frame rate mutations (such as from 30fps to 15fps) are marked, and the mutation frequency is ≤2 times / minute.

[0779] 2. Motion smoothness evaluation:

[0780] Based on the optical flow method, the motion vector between adjacent frames is calculated to evaluate the motion continuity:

[0781] The distribution entropy of the motion vector length is calculated. The lower the entropy value, the smoother the motion.

[0782] For fast motion scenes (such as sports videos), additional motion blur detection is performed to ensure that the blurring degree is ≤0.8px.

[0783] 3. Stutter detection algorithm:

[0784] Detect video playback stall (stutter):

[0785] Based on DCT coefficient change, frame freezing is detected. If it lasts more than 0.5 seconds, it is determined to be stutter.

[0786] Calculate the stutter frequency (times / minute) and stutter cumulative duration, with thresholds of ≤1 times / minute and ≤0.5 seconds / minute, respectively.

[0787] Output smoothness report:

[0788]

[0789]

[0790] The semantic consistency verification module verifies the consistency of the subtitle-picture, audio-visual semantics, and generates a semantic evaluation result, including the following steps:

[0791] 1. Subtitle-picture synchronization detection:

[0792] Based on OCR technology, extract the text in the video picture and compare it with the subtitle text:

[0793] Calculate the edit distance (Levenshtein distance), and the matching degree ≥ 90% is qualified;

[0794] Accurately match the product name, price, and other key information, allowing 0 errors;

[0795] Detect the synchronization of the display time of the subtitle and the content of the picture:

[0796] When a person speaks in the picture, the subtitle should be displayed within 0.3 seconds after the lip movement starts and disappear within 0.2 seconds after it ends.

[0797] 2. Audio-visual semantic alignment:

[0798] Analyze the relevance of audio keywords and video scenes:

[0799] When the audio mentions "left product", the picture should switch to the left product within 3 seconds;

[0800] For emotional expressions (such as excited tone), verify whether the video picture matches (such as fast editing in a promotion scene);

[0801] Build a cross-modal semantic similarity model, calculate the semantic correlation between audio and video frames through the CLIP architecture, and the threshold ≥ 0.7.

[0802] 3. Logical coherence verification:

[0803] Detect the logical coherence of the video content:

[0804] Analyze whether the shot transition conforms to the narrative logic (such as from product demonstration to usage demonstration);

[0805] Verify the matching degree of transition effects and scene changes (such as gradual transition for soothing scenes, fast cut for dynamic scenes);

[0806] Output semantic evaluation report:

[0807]

[0808] Optimization iteration and output module iteratively optimizes and outputs the final video based on the evaluation results, including the following steps:

[0809] 1. Multi-dimensional optimization decision engine:

[0810] Construct a decision tree model to develop optimization strategies based on evaluation results:

[0811] If VMAF < 85, increase the code rate by 15% and re-encode;

[0812] If the frame freezing frequency is > 1 time / minute, adjust the GOP structure (reduce I frame interval);

[0813] If the subtitle matching degree < 0.9, re-perform speech recognition and time axis calibration;

[0814] Support multiple iterations, re-evaluate after each optimization until all thresholds are met.

[0815] 2. Quality-efficiency balancing algorithm:

[0816] Adopt the Pareto optimality principle to find a balance point between quality improvement and computational resource consumption:

[0817] Optimization operations with marginal benefits < 5% in picture quality improvement are automatically terminated;

[0818] Prioritize optimization of user-sensitive areas (such as faces, text), and appropriately reduce quality in secondary areas;

[0819] Support user-defined quality preferences (such as "extreme picture quality" "fast export").

[0820] 3. Final output processing:

[0821] Integrate all optimization results to generate the final video file:

[0822] Add metadata to mark quality indicators (such as VMAF, PSNR);

[0823] Embed quality assessment reports as video attachments;

[0824] Support multiple version outputs (such as different code rates, resolutions) to meet the needs of different distribution channels;

[0825] When outputting the final video, automatically generate thumbnails and preview clips to facilitate content review.

[0826] In summary, picture quality assessment provides technical quality benchmarks, smoothness detection ensures viewing experience, and semantic verification ensures accurate information transmission. If the evaluation finds problems (such as out-of-sync subtitles), the system will backtrack to the relevant module (such as the subtitle positioning optimization sub-unit) for correction, forming a "evaluation → decision → optimization → re-evaluation" closed loop. This process has increased the video quality compliance rate from 78% to 96%, reduced manual review workload by 60%, and reduced the rate of returns due to video quality problems in the e-commerce scene by 23%.

[0827] Figure 2 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. As shown in the figure, an electronic device includes a memory and a processor, the memory stores a computer executable program, and the processor is configured to run the computer executable program to perform the method in any of the above examples. Figure 2

[0828] In the above embodiments, the exemplary explanations about the technical processes of each step can refer to the descriptions above. Figure 2 Figure 1 The above embodiments are only used for describing the embodiments of the present application, but not limit the embodiments of the present application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application should be defined by the claims. The above embodiments of the present application, device, module or unit can be implemented by computer chips or entities, or by products with certain functions.

[0829] For the convenience of description, the above device is described as various units by function. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware when implementing the present application.

[0830] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, the present application, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0831] The present application is described with reference to flowcharts and / or block diagrams of the method, device (the present application), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a machine that implements the flowcharts and / or block diagrams.

[0832] The present application is described with reference to flowcharts and / or block diagrams of the method, device (the present application), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a machine that implements the flowcharts and / or block diagrams. Figure 1 Figure 1 ​​​means for performing the function specified by the block or blocks.

[0833] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a Figure 1 one or more processes and / or blocks Figure 1 means for performing the function specified by the block or blocks.

[0834] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the processes Figure 1 one or more processes and / or blocks Figure 1 means for performing the function specified by the block or blocks.

[0835] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. The memory can include non-persistent memory, random access memory (RAM), and / or non-volatile memory, such as read only memory (ROM) or flash memory, in a computer readable medium. The memory is an example of computer readable media.

[0836] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.

[0837] It is also important to note that the term "comprising" or "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0838] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, an article of manufacture, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-usable program code.

[0839] The present application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0840] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between embodiments refer to each other or replace each other. For the embodiments of the present application, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the part of the method embodiments.

[0841] The above implementation manners are only used to illustrate the embodiments of the present application, but not limit the embodiments of the present application. Those skilled in the technical field related to the present application can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application, and all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application should be defined by the claims. The embodiments of the present application, the device, the module or the unit illustrated in the above embodiments are specifically implemented by a computer chip or an entity, or implemented by a product with certain function.

[0842] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, an article of manufacture, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer-readable program code.

Claims

1. A video mixing and editing system, characterized in that, include: The material preprocessing unit is used to preprocess the original video material to generate standardized editing material; The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions; The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video; The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video.

2. The video mixing and editing system according to claim 1, characterized in that, The material preprocessing unit is used to preprocess the original video material to generate standardized edited material. Specifically, it includes the following steps: The multi-source format standardization subunit is used to perform format-unified conversion on the original video footage to generate format-standardized footage; The spatiotemporal resolution unification subunit is used to normalize the temporal frame rate and spatial resolution of format-standardized materials in order to generate spatiotemporally standardized materials. The content noise filtering subunit is used to perform noise removal and abnormal frame filtering on spatiotemporally normalized material to generate denoised material. The metadata semantic annotation subunit is used to parse and semantically annotate metadata such as shooting parameters and scene information of the denoised material to generate semantically annotated material.

3. The video mixing and editing system according to claim 1, characterized in that, The intelligent editing decision unit is used to perform semantic analysis and editing logic generation on standardized editing materials to generate editing decision instructions, specifically including the following steps: The video semantic understanding subunit is used to identify key content and perform semantic parsing on semantically labeled materials to generate semantic parsing results; The user intent parsing subunit is used to perform natural language understanding and intent extraction processing on the user's input editing requirements in order to generate editing intent instructions; The template matching and strategy generation subunit is used to perform editing template matching and strategy generation processing based on semantic parsing results and editing intent instructions in order to generate editing strategy schemes; The editing logic arrangement subunit is used to perform timeline logic arrangement and keyframe marking processing on the editing strategy scheme in order to generate editing decision instructions.

4. The video mixing and editing system according to claim 1, characterized in that, The automated editing execution unit is used to automatically edit standardized editing materials according to editing decision instructions to generate a preliminary montage video. Specifically, it includes the following steps: The keyframe intelligent extraction subunit is used to perform visual saliency and sentiment peak detection processing on semantically labeled materials to generate a set of keyframes; The shot segmentation and reconstruction subunit is used to perform shot segmentation and temporal reconstruction processing on the keyframe set and editing decision instructions to generate a reconstructed shot sequence; The automatic transition effects generation subunit is used to intelligently match and generate transition effects for the recombined shot sequence in order to generate a shot sequence with transitions. The multitrack audio fusion subunit is used to perform multitrack fusion processing of background sound effects, voice narration and background music on the shot sequence with transitions to generate a preliminary montage video.

5. The video mixing and editing system according to claim 1, characterized in that, The effects optimization output unit is used to enhance the visual effects and optimize the format of the initial mashup video to generate the target mashup video. Specifically, it includes the following steps: The visual effects enhancement subunit is used to perform color correction, sharpening, and special effects overlay processing on the initial montage video to generate a visually enhanced video; The intelligent subtitle generation and positioning subunit is used to perform speech recognition and subtitle generation on visually enhanced videos, and to perform subtitle positioning processing in combination with screen area detection to generate videos with subtitles. The multi-terminal adaptation and optimization subunit is used to perform multi-terminal adaptation and optimization processing on subtitled videos in terms of resolution, bit rate and format, so as to generate adapted and optimized videos. The quality assessment and output subunit is used to evaluate the smoothness, image quality, and semantic consistency of the adapted and optimized video, and to perform optimization processing based on the evaluation results to generate the target mashup video.

Citation Information

Patent Citations

  • Intelligent Mongolian video analysis method

    CN114998785A

  • Short video editing method and system based on artificial intelligence

    CN119031197A

  • Video automatic generation system based on Internet

    CN119299802A

  • Short video clip synthesis method based on fuzzy logic

    CN119854573A

  • High-quality video content automatic generation method and related equipment

    CN120050487A

Cited By

  • Short video editing method and device, electronic equipment, storage medium and program product

    CN121509775A

  • Customizable interactive video production method and system based on AI technology

    CN121908081A