Short video editing method and device, electronic equipment, storage medium and program product
By using structured analysis and intelligent matching technology for director's scripts, we have achieved consistency in shot transitions, audio-visual synchronization, and rhythm control in short video editing. This has solved the problems of logical deviations and stylistic inconsistencies that exist in traditional editing, and improved editing efficiency and quality.
Patent Information
- Application Number
- CN202511713976.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional short video editing suffers from issues such as flawed shot transitions, inconsistent audio-visual styles, and difficulty in maintaining rhythm, leading to high costs and low efficiency.
By structurally analyzing the director's script and using storyboard design information for precise shot editing, combined with semantic understanding of plot keywords and visual language tags, intelligent matching of audio and special effects is achieved, and the editing frames, audio streams and special effects layers are synchronized through a timeline alignment engine.
It significantly improves video editing efficiency and quality, ensures audio-visual synchronization and consistent rhythm control, and reduces production costs and time.
Smart Images

Figure CN121509775A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, specifically to the fields of artificial intelligence technology such as large language models, intelligent agents, and video editing, and particularly to a short video editing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of short video platforms, micro-dramas (short dramas), and AIGC (Artificial Intelligence Generated Content) production tools, users' demand for high-quality, short-cycle, and low-cost video production is growing. Especially in scenarios such as short dramas, advertisements, trailers, and narrative short videos, the traditional manual processes of shot editing, sound mixing, and special effects rendering are not only time-consuming and costly, but also difficult to guarantee in terms of pacing control and stylistic consistency. Summary of the Invention
[0003] This disclosure provides a short video editing method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0004] In a first aspect, embodiments of this disclosure propose a short video editing method, comprising: obtaining a director's script of a target short video to be edited; editing image frames of the target short video based on storyboard design information contained in the director's script; adding matching audio to the target short video based on semantic understanding results of plot keywords contained in the director's script; adding matching special effects to the target short video based on visual language tags and scene tags contained in the director's script; and aligning the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
[0005] Secondly, embodiments of this disclosure propose a short video editing apparatus, comprising: a director's script acquisition unit configured to acquire a director's script of a target short video to be edited; a video editing unit configured to edit image frames of the target short video based on storyboard design information contained in the director's script; an audio adding unit configured to add matching audio to the target short video based on semantic understanding results of plot keywords contained in the director's script; a special effects adding unit configured to add matching special effects to the target short video based on visual language and scene tags contained in the director's script; and a timeline alignment unit configured to align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
[0006] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the short video editing method described in the first aspect.
[0007] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the short video editing method as described in the first aspect when executed.
[0008] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the short video editing method as described in the first aspect.
[0009] The short video editing solution disclosed herein first provides precise shot instructions for image frame editing through storyboard design information, thereby eliminating logical deviations in shot transitions in traditional editing. Secondly, it drives intelligent audio matching through semantic understanding of plot keywords, achieving audio-visual context synchronization through sentiment analysis and environmental analysis. Simultaneously, it guides the dynamic generation of special effects through visual language tags and scene tags to strongly bind them to scene semantics, avoiding stylistic disjointedness. Finally, it synchronizes editing frames, audio streams, and special effects layers along the timeline using a timeline alignment engine, significantly reducing pacing errors in the final product. This solution achieves multimodal collaborative optimization of video editing through structured analysis of the director's script, significantly improving video editing efficiency and the industry standard of the edited results, as well as the production quality of short dramas.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart of a short video editing method provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating an image frame editing method provided in this disclosure embodiment; Figure 4A flowchart illustrating an image frame editing method based on plot rhythm balance provided in this disclosure embodiment; Figure 5 A flowchart illustrating a semantic-based audio addition method provided in this disclosure embodiment; Figure 6 A flowchart illustrating a method for adding matching effects using visual language tags and scene tags, as provided in this embodiment of the disclosure; Figure 7-1 A schematic diagram illustrating the implementation process of adding video subtitles by an editing intelligence agent, provided in an embodiment of this disclosure; Figure 7-2 A schematic diagram illustrating the implementation process of an audio mixing intelligence agent automatically matching sound effects, provided in an embodiment of this disclosure; Figure 8 A structural block diagram of a short video editing device provided in this disclosure embodiment; Figure 9 This is a schematic diagram of the structure of an electronic device suitable for performing a short video editing method, provided as an embodiment of the present disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0013] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0014] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the short video editing methods, apparatuses, electronic devices, and computer-readable storage media of this disclosure can be applied.
[0015] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0016] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include short drama generation applications, video editing applications, and instant messaging applications.
[0017] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0018] Server 105 can provide various services through its built-in applications. Taking a video editing application as an example, when running this application, server 105 can achieve the following: First, it receives the director's script of the target short video to be edited from terminal devices 101, 102, and 103 via network 104; then, based on the storyboard design information contained in the director's script, it edits the image frames of the target short video; simultaneously, based on the semantic understanding results of the plot keywords contained in the director's script, it adds matching audio to the target short video; and based on the visual language tags and scene tags contained in the director's script, it adds matching special effects to the target short video; finally, it aligns the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
[0019] It should be noted that the director's script for the target short video to be edited can be temporarily obtained from terminal devices 101, 102, and 103 via network 104, or it can be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (for example, when starting to process previously stored video editing tasks), it can choose to directly obtain this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0020] Because video editing requires significant computing resources and power, the short video editing methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the short video editing device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through video editing applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, if the video editing application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the short video editing device can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0022] Please refer to Figure 2 , Figure 2 A flowchart of a short video editing method provided in this disclosure embodiment, wherein process 200 includes the following steps: Step 201: Obtain the director's script for the target short video to be edited; This step is intended for the implementer of short video editing methods (e.g., Figure 1 The server 105 shown retrieves the director's script for the target short video to be edited. This target short video can be a segment of a short drama, an advertisement, a game trailer, or other content from a short video platform. The director's script is a collection of metadata containing timecode, shot sequences, visual instructions, and narrative logic. Specifically, the executing entity can obtain key information that represents the editing intent by structured analysis of the director's script: shot design information (such as shot start and end timestamps, shot classification labels), visual language labels (such as "warm and cool tones" and "dynamic camera movement"), and scene and plot semantics (such as the emotional intensity value of "confrontation in the rainy night").
[0023] Step 202: Based on the storyboard design information contained in the director's script, edit the image frames of the target short video; Building upon step 201, this step aims to have the aforementioned executing entity edit the image frames of the target short video based on the storyboard design information contained in the director's script. The storyboard design information in the director's script includes shot planning (such as splicing order and transition methods) and character introduction planning (such as the timing of introducing new characters).
[0024] Specifically, the process begins by extracting shot planning and character entrance planning using a structured analysis engine. A temporal relationship analysis algorithm is then employed to identify shot sequences and transition markers. An entity recognition model is used to locate the keyframe positions where a new character first appears. Finally, the original video frames corresponding to each shot are arranged sequentially. Optical flow analysis can also be used to detect the continuity of motion trajectories between adjacent shots. For example, if a break in motion is detected (such as switching to shot 2 before the waving motion in shot 1 is completed), automatic interpolation is used to generate transition frames (inserting 3 frames to smooth the arm trajectory). Simultaneously, a storyboard association database is established. When a subsequent storyboard references a previous scene (such as a flashback), the lighting parameters and composition ratios of the previous storyboard are forcibly inherited to eliminate any sense of disjointed editing.
[0025] Step 203: Based on the semantic understanding results of the plot keywords contained in the director's script, add matching audio to the target short video; Building upon step 201, this step aims to have the aforementioned executing entity add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script.
[0026] Specifically, the execution entity can first extract the emotional tone, identify environmental features, and dynamic events from the director's script; then, it performs a multimodal search in a preset sound effects library, matching existing audio resources using acoustic feature vectors (such as low-frequency energy distribution and pulse density). If no direct match is found in the library, basic sound effects are generated based on a physical simulator; emotional features are then injected through an emotion enhancement module; finally, after decomposing the dialogue into phoneme sequences, the speech rate (increased by 40% for angry speech) and pitch fluctuations (decreased by 15% for sad scenes) are adjusted according to the emotional value. Furthermore, environmental changes can be simulated using an acoustic propagation model. For example, when a scene transition between adjacent shots is detected (such as "indoor → outdoor"), the difference in reverberation parameters is automatically calculated to generate a gradual audio transition.
[0027] In this embodiment, the executing entity can extract voiceprint features when a character's lines first appear, establish a character ID-voiceprint mapping library, and ensure the consistency of the voice features of the same character in subsequent shots.
[0028] Step 204: Based on the visual language tags and scene tags contained in the director's script, add matching special effects to the target short video; Building upon step 201, this step aims to have the aforementioned executing entity add matching special effects to the target short video based on the visual language tags and scene tags contained in the director's script. The addition of matching special effects based on the visual language tags and scene tags in the director's script is achieved through a semantically driven special effects generation mechanism. Visual language tags can define basic stylistic attributes of the image, such as "cool color tone" or "motion blur," while scene tags are used to trigger environmental-level special effects combinations, such as "rain particles + ground reflection / holographic projection + laser beam."
[0029] In this embodiment, the executing entity can identify the type of special effects through a label classification model, then locate the effective range of the special effects based on timestamps, and finally perform layered compositing through the rendering engine. Specifically, the base layer adjusts the global color curve, the overlay layer injects dynamic elements, and physical simulation is applied to enhance realism. Furthermore, when label correlation is detected in consecutive scenes, gradient filters can be automatically generated to avoid visual jumps. For example, if the plot rhythm label shows "climax conflict," the special effects' expressiveness is enhanced; if it's "calm dialogue," the particle density is reduced to the base value.
[0030] Step 205: Align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
[0031] Building upon steps 202-204, this step aims to align the edited image frames, the added audio stream, and special effects along the timeline, resulting in the edited target short video. Specifically, the executing entity can first establish a main timeline based on the director's script timecode, with each material element bound to an independent timestamp: the image frame sequence inherits the start time of the storyboard, the audio stream is anchored according to dialogue nodes, and the special effects layer is associated with scene transition times; then, through a dynamic delay compensation algorithm, the system detects the processing delay of different modal data in real time and automatically adjusts the audio track pre-compensation to ensure strict alignment of lip movements and pronunciation.
[0032] The rendering process employs a layered rendering strategy to align the edited image frames, added audio streams, and effects along the timeline. Specifically, the bottom layer arranges the video tracks in frame sequence order, the middle layer overlays the effects particle system, and the top layer mounts the audio waveform. When multi-track overlap is detected, the rendering process activates a conflict resolution mechanism, prioritizing core narrative elements (dialogue audio has absolute priority) and dynamically compressing the duration of non-critical effects. During mobile rendering, particle density is automatically simplified, and keyframe thinning technology maintains smoothness on low-power devices.
[0033] Furthermore, when the audio stream is unexpectedly interrupted, supplementary speech can be generated based on contextual semantics; when special effects rendering times out, a backup template can be enabled (e.g., replacing complex particles with simplified light effects).
[0034] The short video editing method provided in this disclosure firstly provides precise shot instructions for image frame editing through storyboard design information, thereby eliminating logical deviations in shot transitions in traditional editing. Secondly, it drives intelligent audio matching through semantic understanding of plot keywords, achieving audio-visual context synchronization through sentiment analysis and environmental analysis. Simultaneously, it guides the dynamic generation of special effects through visual language tags and scene tags to strongly bind them to scene semantics, avoiding stylistic disjointedness. Finally, it uses a timeline alignment engine to synchronize editing frames, audio streams, and special effects layers along the timeline, significantly reducing rhythm errors in the final product. This solution achieves multimodal collaborative optimization of video editing through structured analysis of the director's script, significantly improving video editing efficiency and the industry standard of the edited results, as well as the production quality of short dramas.
[0035] To accurately map the storyboard descriptions to the physical editing actions, please refer to... Figure 3 , Figure 3 A flowchart of an image frame editing method provided for an embodiment of this disclosure, wherein process 300 includes the following steps: Step 301: Determine the splicing order and transition method of the image frames of each shot that constitutes the target short video based on the shot planning contained in the director's script; This step aims to have the aforementioned executing entity determine the splicing order and transition methods of the image frames constituting the target short video based on the shot planning contained in the director's script. Specifically, the shot planning contained in the director's script can be parsed, a timeline mapping relationship can be established using the shot transition points as anchor points, the splicing order of the image frames of each shot can be determined through a temporal relationship analysis algorithm, and the scene attributes of adjacent shots can be detected. If the difference in scene type exceeds a threshold, the transition instruction generation process is triggered to generate the transition method.
[0036] Step 302: Stitch the image frames of each shot in the stitching order, and add a transition frame between the last frame and the first frame of two consecutive shots with scene changes after stitching. In this embodiment, the execution entity stitches the image frames of each shot in the stitching order, and adds a transition frame between the last frame and the first frame of two consecutive shots that have a scene change after stitching, according to the transition method. Specifically, when a shot pair that needs a transition is detected (such as the last frame of shot A and the first frame of shot B), the execution entity can generate a transition frame according to the transition method specified in the script.
[0037] Step 303: Based on the character entrance plan contained in the director's script, generate a subtitle bar describing the new character for the first image frame in the target short video that presents the new character.
[0038] Furthermore, the executing entity can also generate subtitles describing the newly introduced character based on the character entrance plan included in the director's script, for the first image frame in the target short video that presents the new character. The generated subtitles have a style that matches the style of the first image frame presenting the new character and / or the style of the new character, and are placed in a position within the first image frame that does not overlap with any other subject.
[0039] When a new character's appearance frame is detected, the main subject area is defined through face detection. A safe zone is calculated using compositional aesthetic rules. Style matching is then used to determine a style that matches the style of the first image frame presenting the new character and / or the style of the new character. Corresponding subtitles are then generated within the safe zone according to this style. The style of the subtitles can be determined by extracting the main color tone and texture features of the image. Subsequent appearances of the same character automatically inherit the initially determined subtitle style, ensuring visual consistency.
[0040] Furthermore, users can manually adjust transition durations or subtitle positions during editing. In addition, when new characters appear consecutively in the same scene, the subtitle background transparency and animation are automatically standardized to avoid visual interference.
[0041] The image frame editing method provided in this embodiment first determines the splicing order and transition methods of the image frames constituting the target short video based on the shot planning contained in the director's script; then, based on the character entrance planning contained in the director's script, it generates a subtitle bar describing the newly introduced character for the first image frame in the target short video; finally, based on the character entrance planning contained in the director's script, it generates a subtitle bar describing the newly introduced character for the first image frame in the target short video. This embodiment determines the shot sequence and transition methods by deeply analyzing the structured data of the director's script, ensuring the smoothness of the subsequently generated video, and transforming the error-prone manual annotation and editing process into a high-precision, reusable automated workflow, significantly shortening the video production cycle.
[0042] To achieve precise optimization of video narrative pacing based on plot rhythm, please refer to... Figure 4 , Figure 4 A flowchart of an image frame editing method based on plot rhythm balance provided in this disclosure embodiment, wherein process 400 includes the following steps: Step 401: Determine the plot rhythm information based on the emotional changes of the characters in the director's script; In this embodiment, the executing entity determines the plot rhythm information based on the emotional change information of the characters appearing in the director's script. Specifically, the emotional change information of the characters in the director's script can be analyzed first, and the text description can be converted into a numerical sequence to generate an emotional curve through an emotional intensity quantification model. Then, the emotional curve can be scanned in time units of a preset duration to generate the plot rhythm information.
[0043] Step 402: Generate plot rhythm tags for the target short video based on the plot rhythm information; Building upon step 401, this step aims to have the aforementioned executing entity generate plot rhythm tags for the target short video based on the plot rhythm information. Specifically, when an emotional abrupt change is detected, the image frame corresponding to the emotional abrupt change in the target short video can be marked as a rhythm keyframe, and plot rhythm tags can be automatically generated. For example, a "conflict peak" can be marked for a climax, and a "buffer zone" can be marked for a transitional section.
[0044] Step 403: Calculate the rhythm index based on the video duration in the middle of different plot rhythm tags, and adjust the rhythm of the target short video to the desired balanced rhythm based on the rhythm index.
[0045] Building upon step 402, this step aims to have the aforementioned executing entity calculate a rhythm index based on the video duration within different plot rhythm tags, and adjust the rhythm of the target short video to the desired balanced rhythm based on the rhythm index. Specifically, the rhythm index can be calculated by statistically analyzing the cumulative duration percentage corresponding to each plot rhythm tag. This rhythm index can include the amplitude of emotional fluctuations and the frequency of segment transitions. When the rhythm index exceeds a preset threshold, the plot rhythm is determined to be unbalanced.
[0046] Adjusting the pace of a target short video to a desired balanced pace based on the pace index can include the following methods: For segments with an overly fast pace, dynamic blur frames can be inserted to prolong visual dwell time; for segments with a sluggish pace, an intelligent editing engine can be enabled to automatically remove redundant actions; when the emotional lines of multiple characters intertwine, a composite pace curve can be generated through an emotion overlay algorithm to force the peak of the main character's emotion to align with the shot transition point; and models can also be trained based on historical data to automatically insert environmental empty shots in emotionally depressing segments to avoid viewer fatigue.
[0047] This embodiment discloses an image frame editing method based on plot rhythm balance. First, it determines plot rhythm information based on the emotional changes of the characters in the director's script. Then, it generates plot rhythm tags for the target short video based on this information. Finally, it calculates a rhythm index based on the video duration between different plot rhythm tags and adjusts the target short video's rhythm to the desired balanced rhythm. This embodiment, through emotion-driven rhythm generation, transforms character psychological changes into a visual narrative rhythm, ensuring that the video rhythm always approaches the optimal experience threshold. This solves the problem of the disconnect between emotional expression and rhythm control, achieving intelligent optimization of plot rhythm.
[0048] Building upon the aforementioned implementation, to optimize for redundant and abnormal frames, the execution entity can first identify redundant image frames and / or abnormal frames in the target short video; then, it can remove the identified redundant image frames; finally, it can repair the identified abnormal frames according to their abnormality type using appropriate repair methods. The abnormality types include: still frames, blurred frames, and dialogue misalignment frames. For redundant frame detection, the execution entity uses a time-sliding window algorithm to calculate the visual similarity of consecutive frames; when more than 5 consecutive frames meet the condition, it is determined to be a redundant sequence. For abnormal frame detection, the execution entity establishes a categorized recognition model: still frames are captured by using a pixel change rate threshold to detect image stillness; blurred frames are calculated using a Laplacian operator to calculate image sharpness; and dialogue misalignment frames are detected by an audio-visual synchronization analyzer to detect lip movements and audio timestamp deviations.
[0049] Based on the above implementation, the execution entity can perform intelligent deletion of redundant frames in the following ways: retain key action node frames, generate intermediate transition frames through motion trajectory interpolation algorithms to ensure action continuity; for still frames, an optical flow prediction model can be called to generate a compensation frame sequence based on the motion vectors of the preceding and following frames; for blurred frames, a Generative Adversarial Network (GAN) super-resolution reconstruction network can be used to improve detail clarity, while constraining texture generation to conform to scene consistency; for dialogue misalignment frames, a two-way adjustment mechanism can be adopted, if the video is lagging, the frame sequence is moved forward, if the audio is lagging, the silent segment is dynamically cropped, and the lip-sync animation is fine-tuned through a phoneme alignment model.
[0050] Building upon the aforementioned implementation, the executing entity can also automatically record frequently occurring anomaly types and preload optimization parameters for subsequent similar scenarios. Simultaneously, the executing entity can establish a cross-segment repair knowledge base, triggering global parameter optimization when a similar defect pattern is detected.
[0051] To achieve a high degree of integration between audio and video, please refer to... Figure 5 , Figure 5A flowchart of a semantic-based audio addition method provided for embodiments of this disclosure, wherein process 500 includes the following steps: Step 501: Determine the semantics of the plot keywords extracted from the director's script; In this embodiment, the executing entity determines the semantics of the plot keywords extracted from the director's script. Specifically, the plot keywords are first extracted from the director's script, then semantic parsing is performed using a pre-trained natural language processing model, and word vector encoding technology is used to convert the text into semantic vectors, which serve as the semantics of the plot keywords.
[0052] In this embodiment, the executing entity can construct a multi-level semantic index. The first-level index is associated with the basic scene type (e.g., nature, city, science fiction, etc.), the second-level index is bound to dynamic parameters (e.g., rain sound intensity 0-1 scale), and the third-level index is linked to spatiotemporal attributes (e.g., cave echo attenuation coefficient).
[0053] Step 502: Determine existing background sounds and / or existing ambient sounds that match the semantics from the preset sound effects library; In this embodiment, the executing entity can determine existing background sounds and / or existing ambient sounds that match the semantics from a preset sound effects library. Specifically, the executing entity compares the semantic vector with the audio feature vectors of sound effects in the preset sound effects library using cosine similarity. When the similarity is greater than a preset threshold, the current sound effect is determined to be an existing background sound and / or existing ambient sound that matches the semantics.
[0054] Step 503: In response to the absence of a semantically matching background sound and / or ambient sound in the sound effects library, generate a new semantically matching background sound and / or ambient sound using a preset audio generation model; In this embodiment, when there is no background sound and / or ambient sound in the sound effect library that matches the semantics, the execution subject inputs the semantic vector and physical parameter constraints into the preset audio generation model to generate a new background sound and / or a new ambient sound that matches the semantics.
[0055] Step 504: Add the semantically matched audio to the video frame.
[0056] In this embodiment, the executing entity adds semantically matching audio to the video frame corresponding to the relevant plot keyword. Specifically, the executing entity inserts semantically matching audio into the audio track of the video frame corresponding to the relevant plot keyword, based on the timecode and keyframes in the director's script.
[0057] In this embodiment, when multiple audio frequencies are detected to overlap, a priority strategy can be automatically activated to adjust the multiple audio frequencies. For example, when dialogue and explosion sounds overlap, the sound pressure level of the dialogue is increased, and the amplitude of the explosion sounds is reduced. The executing entity can also adjust the volume curve of the rain sound in real time according to the visual elements on the screen. For example, when the visual element on the screen is raindrops, the volume curve of the rain sound is adjusted in real time according to the density of the raindrops.
[0058] In this embodiment, to ensure optimal performance of each audio element in different plot scenarios, the executing entity can determine a mixing balance strategy for the dialogue audio, background sound, and ambient sound in the target short video, and then perform mixing processing on the dialogue audio, background sound, and ambient sound according to the mixing balance strategy. The specific mixing processing method is as follows: First, the fundamental frequency range of the dialogue audio is extracted, the energy distribution of the main melody of the background music is identified, and the intensity of the ambient sound's continuous noise is determined. Then, an adaptive masking algorithm is used to automatically reduce the gain of the corresponding frequency band of the background music by calculating the spectral overlap between the dialogue and the background sound, and adjust the ambient sound according to the noise intensity to avoid acoustic masking effects that reduce the clarity of the dialogue. For example, when the plot tag is "intense conflict," the maximum loudness of the ambient sound is limited to -6dBFS, the dialogue sound pressure level is increased by 15%, and the dynamic range is expanded so that special effects such as explosions do not drown out key lines; in "lyrical passages," the high-frequency overtones of the background music can be increased, while the ambient sound is faded in and out to enhance the immersive atmosphere. In addition, when adjacent shots switch scenes, the system automatically inherits the noise floor of the previous shot (such as air conditioner noise 30dB) and superimposes the ambient sound of the new scene (traffic noise 40dB), eliminating auditory jumps through linear gradation.
[0059] The semantic-based audio addition method disclosed in this embodiment first determines the semantics of plot keywords extracted from the director's script. Then, it identifies existing background sounds and / or ambient sounds that match the semantics from a preset sound effects library. When no matching background sounds and / or ambient sounds exist in the sound effects library, a new matching background sound and / or ambient sound is generated using a preset audio generation model. Finally, the semantically matching audio is added to the video frame. This embodiment, by deeply binding plot semantics with auditory experience, transforms sound effects design from being dominated by human experience to data-driven automated production, significantly improving the emotional immersion and production efficiency of video content. Furthermore, it generates customized audio when no matching sound effects are found in the sound effects library, ensuring audio-visual consistency.
[0060] Furthermore, in mobile playback scenarios, the frequency response defects of mobile phone speakers can be compensated by enhancing the dialogue frequency band; while for headphone users, the binaural rendering algorithm can be activated to enhance the spatial orientation of ambient sounds.
[0061] To achieve a deep integration of visual effects and story scenes, please refer to... Figure 6 , Figure 6 A flowchart of a method for adding matching effects using visual language tags and scene tags provided in this disclosure embodiment is provided, wherein process 600 includes the following steps: Step 601: Generate text animations, character effects, background effects, and filter effects based on visual language tags; In this embodiment, the executing entity generates text animations, character effects, background effects, and filter effects based on visual language tags. Specifically, the executing entity parses the visual language tags to generate text animations, character effects, background effects, and filter effects. Text animations refer to generating text animation trajectories using a path algorithm; for example, the keyword "unveil" triggers a text fragmentation and reconstruction animation. Character effects refer to adding dynamic elements based on a skeletal binding model; for example, the tag "magician" triggers fingertip particle streams. Background effects refer to simulating environmental interactions through a physics engine; for example, the tag "rainstorm" generates ripples spreading as raindrops collide with the ground. Filter effects refer to applying a color transfer model; for example, the tag "nostalgia" overlays sepia tones and grainy noise.
[0062] Step 602: Determine the scene change time based on the scene tags, and generate highlight visual effects for the scene change time; In this embodiment, the executing entity determines the scene change time based on the scene label and generates a highlight visual effect for the scene change time. Specifically, the executing entity determines the scene change time based on the scene label, generates a highlight effect for the scene change time, and generates a transition effect sequence before and after the switch.
[0063] Step 603: Determine the plot rhythm based on the scene tags, and control the style of each special effect within each plot rhythm to match the corresponding plot rhythm.
[0064] In this embodiment, the executing entity determines the plot rhythm based on scene tags and controls the style of various special effects within each plot rhythm to match the corresponding plot rhythm. For example, when a "intense conflict" rhythm is identified based on the scene tag, the particle effect emissivity is increased and the filter contrast is enhanced; when a "lyrical passage" rhythm is identified based on the scene tag, a low-interference mode is activated, reducing the background particle density. In addition, the same plot passage can inherit the main color tone and motion blur parameters, and deviations are automatically corrected through optical flow consistency detection.
[0065] In this embodiment, the execution entity can achieve device-adaptive rendering, such as automatically simplifying the particle system on mobile devices and maintaining performance through keyframe thinning. Furthermore, when a conflict between special effects and scene semantics is detected, special effects correction operations can be performed; for example, when "raindrops" appear in a "desert" scene, the system automatically switches to a sandstorm particle model.
[0066] This embodiment discloses a method for adding matching special effects using visual language tags and scene tags. First, it generates text animations, character effects, background effects, and filter effects based on the visual language tags. Then, it determines the scene change time based on the scene tags and generates highlight visual effects for that time. Finally, it determines the plot rhythm based on the scene tags and controls the style of each effect within each plot rhythm to match the corresponding rhythm. This embodiment significantly lowers the barrier to special effects production through large-scale automated production of effects. By using tags as a medium to concretize abstract plots into effect parameters, it achieves a precise visual transformation of narrative intent, providing a fully automated and scalable special effects solution for content creation.
[0067] Based on the above implementation, the executing entity can also utilize multiple intelligent agents to add special effects to the target short video. For example, the executing entity can first use a preset video editing intelligent agent to edit the image frames of the target short video based on the storyboard design information contained in the director's script; then use a preset audio adding intelligent agent to add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script; finally, use a preset special effects adding intelligent agent to add matching special effects to the target short video based on the visual language and scene tags contained in the director's script.
[0068] Building upon the aforementioned implementation, the executing entity can further conduct a multi-dimensional quality assessment of the edited target short video and generate modification suggestions if the multi-dimensional quality assessment fails. Based on these suggestions, adjustments can be made to at least one of the video images, audio, and special effects until the adjusted target short video passes the multi-dimensional quality assessment. The quality dimensions assessed in the multi-dimensional quality assessment include at least two of the following: frame continuity and image consistency when connecting two consecutive shots, audio-visual synchronization, sound field equalization, composition score, color distribution score, and rhythmic integrity.
[0069] Specifically, the implementing entity can establish the following multi-dimensional quality assessment system: 1) Use optical flow analysis to detect the continuity of image frames, and identify smoothness issues in shot transitions by calculating the mutation rate of pixel motion vectors between adjacent image frames; 2) Evaluate the consistency of the scene based on feature point matching algorithms, and trigger a scene transition alarm if the matching degree of key feature points in consecutive shots is lower than a preset threshold; 3) Detect audio-visual synchronization by aligning lip movements and phoneme timestamps at the millisecond level; 4) Analyze the spectral energy distribution of dialogue, background sounds, and ambient sounds using sound field equalization, for example, if the 200-4000Hz frequency band of dialogue is obscured by more than 30% by background music, it is considered unbalanced; 5) Calculate the center-of-gravity shift of the scene based on the rule of thirds and the subject focus algorithm to score the composition; 6) Score the color distribution by comparing the main color difference values of adjacent shots using HSV (hue, saturation, brightness) space histograms; 7) Combine the emotional curve of the plot with the shot duration distribution to judge the completeness of the rhythm and detect whether the climax segment reaches the preset duration percentage.
[0070] Building upon the aforementioned implementation, when a multi-dimensional quality assessment fails, the executing entity activates an intelligent diagnostic engine to pinpoint the root cause and generate suggested modifications. For example, upon detecting a frame consistency defect, it automatically marks the range of problematic frames and suggests adding a 0.3-second transition effect; when sound field imbalance is detected, it outputs an audio correction parameter of "4dB attenuation in the mid-frequency band of background music." Furthermore, the executing entity can employ a layered repair strategy: for frame continuity issues, it calls a frame interpolation algorithm to generate a transition sequence; when compositional defects exist, it triggers an automatic cropping and dynamic reconstruction module; and when rhythm is missing, it inserts empty shots or compresses redundant segments based on an emotional intensity model.
[0071] Building upon the above implementation, when audio-visual asynchrony and abnormal color distribution are detected simultaneously, audio repair may affect the timeline; therefore, the audio-visual synchronization issue is repaired first, followed by color adjustment. The execution entity can also record frequently occurring problem scenarios, allowing for pre-loading of optimization parameters for subsequent similar scenarios.
[0072] Based on the above implementation, the executing entity can also specifically use a preset editing evaluation agent to conduct a multi-dimensional quality assessment of the edited target short video, and generate modification suggestions when the multi-dimensional quality assessment fails. Based on the modification suggestions, at least one of the video images, audio and special effects is adjusted until the adjusted target short video passes the multi-dimensional quality assessment.
[0073] To deepen understanding, Figure 7-1 This is a schematic diagram illustrating the implementation process of adding video subtitles using an editing intelligence agent, as provided in an embodiment of this disclosure. Figure 7-2 This is a schematic diagram illustrating the implementation process of an audio mixing intelligence agent automatically matching sound effects, as provided in this embodiment of the disclosure. This disclosure combines... Figure 7-1 and Figure 7-2The specific application scenario shown provides a concrete implementation scheme, as follows: This implementation constructs a multi-agent collaborative system for the automated production of short drama videos. The system includes four main functional agents: a video editing agent, an audio mixing agent, a special effects editing agent, and an editing evaluation agent. They work together to complete the entire process control from shot splicing to special effects rendering and editing evaluation.
[0074] I. Video Editing Intelligent Agent
[0075] The video editing agent is the core execution unit in the system, capable of automatically completing the following tasks based on shot planning, character introductions, and emotional curves in the director's script: 1. Utilize multimodal structured metadata and visual understanding models to achieve automatic planning of shot splicing order and control of transition rhythm. 2. Detect redundant and abnormal frames (such as still frames, blurry frames, misaligned dialogue, etc.) and trigger the "intelligent repair" mechanism to enhance the smoothness of the screen.
[0076] 3. Generate editing rhythm templates based on plot rhythm tags, and achieve rhythm balance through rhythm index calculation and adjustment.
[0077] 4. Automatically generate subtitles that conform to the semantics of character appearance and scene changes, match styles and animations, and perform spatial calibration to ensure that the subtitles do not obscure the main subject and maintain a consistent style.
[0078] II. Audio Mixing Intelligent Agent
[0079] The audio mixing agent is responsible for the selection and mixing control of dialogue, background music, and ambient sound, and has the following capabilities: 1. Based on script keywords, perform semantic extraction and embedding (vector) retrieval to select the most suitable background music and ambient sounds from the sound effects library.
[0080] 2. When a suitable audio source is lacking, use an AI-generated model to synthesize music segments or sound effects.
[0081] 3. Utilize the RAG (Retrieval-Augmented Generation) mechanism to construct cue word templates, guide the mixing balance strategy, and achieve coordination of spectral energy and sound pressure level between dialogue, ambient sound, and background music.
[0082] III. Special Effects Editing Intelligent Agent
[0083] The special effects editing AI automatically matches and generates corresponding visual effects based on the visual language and scene tags marked in the director's script, including: 1. The pacing of the plot is compared with a professional camera knowledge base to ensure that the special effects style is synchronized with the plot.
[0084] 2. Automatically identify key action frames and insert special effects characters, backgrounds, or filters.
[0085] 3. Generate text animations and highlight visual effects at key plot points to enhance narrative expressiveness and visual appeal.
[0086] IV. Editing Evaluation Agent
[0087] To enable the system's self-learning optimization capabilities, this implementation proposes an editing evaluation agent for multi-dimensional evaluation of the final output. Its specific functions are as follows: 1. Detect the inter-frame continuity and image consistency of shot transitions to quantify smoothness.
[0088] 2. Evaluate the degree of audio-visual synchronization and sound field balance.
[0089] 3. Scoring and analyzing the integrity of composition, color, and rhythm to form quantitative aesthetic indicators.
[0090] The system feeds back multi-dimensional evaluation results to each execution sub-agent, constructing an optimization loop for editing quality. Furthermore, the multi-dimensional evaluation supports automated iterative learning, continuously adjusting transition rhythms, editing beats, and audio-visual fusion strategies in subsequent editing, significantly improving the professional quality of the automatically generated final product.
[0091] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a short video editing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0092] like Figure 8As shown, the short video editing device 800 of this embodiment may include: a director script acquisition unit 801, a video editing unit 802, an audio adding unit 803, a special effects adding unit 804, and a timeline alignment unit 805. The director script acquisition unit 801 is configured to acquire the director script of the target short video to be edited; the video editing unit 802 is configured to edit the image frames of the target short video based on the storyboard design information contained in the director script; the audio adding unit 803 is configured to add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director script; the special effects adding unit 804 is configured to add matching special effects to the target short video based on the visual language and scene tags contained in the director script; and the timeline alignment unit 805 is configured to align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
[0093] In this embodiment, the specific processing and technical effects of the following components in the short video editing device 800—namely, the director script acquisition unit 801, the video editing unit 802, the audio addition unit 803, the special effects addition unit 804, and the timeline alignment unit 805—can be found in reference to [the relevant documentation]. Figure 2 The relevant descriptions of steps 201-205 in the corresponding embodiments will not be repeated here.
[0094] In some other implementations of this embodiment, the video editing unit 802 includes: a splicing order and transition method determination subunit, configured to determine the splicing order and transition method of the image frames of each shot constituting the target short video according to the shot planning contained in the director's script; and a splicing and transition frame addition subunit, configured to splice the image frames of each shot in the splicing order, and add a transition frame between the last frame and the first frame of two consecutive shots with scene transition after splicing, according to the transition method.
[0095] In some other implementations of this embodiment, the video editing unit 802 further includes a new character subtitle generation subunit, which is configured to generate a subtitle describing the new character for the first image frame in the target short video that presents the new character, based on the character appearance plan contained in the director's script.
[0096] In some other implementations of this embodiment, the generated subtitle bar has a style that matches the style of the first image frame in which the new character appears and / or the style of the new character, and the subtitle bar is placed in a position in the first image frame in which the new character appears without overlapping with any subject.
[0097] In some other implementations of this embodiment, the video editing unit 802 further includes: a plot rhythm information determination subunit, configured to determine plot rhythm information based on the emotional change information of the characters appearing in the director's script; a plot rhythm tag generation subunit, configured to generate plot rhythm tags for the target short video based on the plot rhythm information; and a plot rhythm balancing subunit, configured to calculate a rhythm index based on the video duration between different plot rhythm tags, and adjust the rhythm of the target short video to the desired balanced rhythm based on the rhythm index.
[0098] In some other implementations of this embodiment, the video editing unit 802 further includes: a redundant frame / abnormal frame identification subunit, configured to identify redundant image frames and / or abnormal frames in the target short video; a redundant frame processing subunit, configured to delete the identified redundant image frames; and an abnormal frame processing subunit, configured to repair the identified abnormal frames according to the abnormal type using the corresponding repair method; wherein, the abnormal types include: still frames, blurred frames, and dialogue misalignment frames.
[0099] In some other implementations of this embodiment, the audio adding unit 803 includes: a semantic determination subunit configured to determine the semantics of plot keywords extracted from the director's script; an existing audio matching subunit configured to determine existing background sounds and / or existing ambient sounds that match the semantics from a preset sound effects library; a new audio generation subunit configured to generate new background sounds and / or new ambient sounds that match the semantics using a preset audio generation model in response to the absence of background sounds and / or ambient sounds that match the semantics in the sound effects library; and an audio adding subunit configured to add the audio that matches the semantics to the video frame to which the corresponding plot keyword belongs.
[0100] In some other implementations of this embodiment, the audio adding unit 803 further includes: a mixing balance strategy determination subunit, configured to determine a mixing balance strategy for dialogue audio, background sound and ambient sound in the target short video; and a mixing balance processing subunit, configured to perform mixing processing on the dialogue audio, background sound and ambient sound according to the mixing balance strategy.
[0101] In some other implementations of this embodiment, the special effects adding unit 804 includes: a first special effects generation subunit, configured to generate text animation, character special effects, background special effects and filter special effects according to visual language tags; and a second special effects generation subunit, configured to determine the scene change time according to scene tags and generate highlight visual effects for the scene change time.
[0102] In some other implementations of this embodiment, the special effects adding unit 804 further includes: a special effects style control unit, configured to determine the plot rhythm according to the scene tag, and control the style of each special effect within each plot rhythm to match the plot rhythm.
[0103] In some other implementations of this embodiment, the video editing unit 802 is further configured to: use a preset video editing agent to edit the image frames of the target short video based on the storyboard design information contained in the director's script; the audio adding unit 803 is further configured to: use a preset audio adding agent to add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script; and the special effects adding unit 804 is further configured to: use a preset special effects adding agent to add matching special effects to the target short video based on the visual language and scene tags contained in the director's script.
[0104] In some other implementations of this embodiment, the short video editing device 800 further includes: a multi-dimensional quality assessment unit configured to perform a multi-dimensional quality assessment on the edited target short video and generate modification suggestions when the multi-dimensional quality assessment fails; and a modification suggestion adjustment unit configured to adjust at least one of the video image, audio, and special effects according to the modification suggestions until the adjusted target short video passes the multi-dimensional quality assessment.
[0105] In some other implementations of this embodiment, the multi-dimensional quality assessment unit and the adjustment unit based on modification opinions are further configured to: perform multi-dimensional quality assessment on the edited target short video using a preset editing evaluation agent, generate modification opinions when the multi-dimensional quality assessment fails, and adjust at least one of the video image, audio and special effects according to the modification opinions until the adjusted target short video passes the multi-dimensional quality assessment.
[0106] In some other implementations of this embodiment, the quality dimensions evaluated by the multidimensional quality assessment include at least two of the following: inter-frame continuity and image consistency when image frames of two consecutive shots are joined, audio-visual synchronization, sound field equalization, composition score, color distribution score, and rhythmic integrity.
[0107] This embodiment exists as a device embodiment corresponding to the above method embodiment. The short video editing device provided in this embodiment first provides precise shot instructions for image frame editing through storyboard design information, thereby eliminating logical deviations in shot transitions in traditional editing. Second, it drives intelligent audio matching through semantic understanding of plot keywords, that is, it achieves audio-visual context synchronization through sentiment analysis and environmental analysis. At the same time, it guides the dynamic generation of special effects through visual language tags and scene tags to strongly bind with scene semantics and avoid style fragmentation. Finally, it synchronizes the editing frames, audio stream, and special effects layers on the timeline through a timeline alignment engine, significantly reducing the rhythm error of the final product. This solution achieves multimodal collaborative optimization of video editing through structured analysis of the director's script, which can not only significantly improve the efficiency of video editing and the industry standard of the edited results, but also significantly improve the production level of short dramas.
[0108] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the short video editing method described in any of the above embodiments.
[0109] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the short video editing method described in any of the above embodiments when executed.
[0110] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the short video editing method described in any of the above embodiments.
[0111] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0112] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded into random access memory (RAM) 903 from storage unit 908. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0113] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0114] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the short video editing method. For example, in some embodiments, the short video editing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the short video editing method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the short video editing method by any other suitable means (e.g., by means of firmware).
[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0120] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0121] According to the technical solution of this disclosure, firstly, precise shot instructions are provided for image frame editing through storyboard design information, thereby eliminating logical deviations in shot transitions in traditional editing. Secondly, semantic understanding of plot keywords drives intelligent audio matching, that is, audio-visual context synchronization is achieved through sentiment analysis and environmental analysis. Simultaneously, visual language tags and scene tags guide the dynamic generation of special effects to strongly bind them with scene semantics, avoiding stylistic disjointedness. Finally, a timeline alignment engine synchronizes the editing frames, audio stream, and special effects layers on the timeline, significantly reducing rhythm errors in the final cut. This solution achieves multimodal collaborative optimization of video editing through structured analysis of the director's script, which can not only significantly improve video editing efficiency and the industry standard of the edited results, but also significantly improve the production quality of short dramas.
[0122] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0123] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A short video editing method, comprising: Obtain the director's script for the target short video to be edited; Based on the storyboard design information contained in the director's script, the image frames of the target short video are edited; Based on the semantic understanding results of the plot keywords contained in the director's script, matching audio is added to the target short video; Based on the visual language tags and scene tags contained in the director's script, add matching special effects to the target short video; Align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited target short video.
2. The method according to claim 1, wherein, The step of editing image frames of the target short video based on the storyboard design information contained in the director's script includes: The stitching order and transition method of the image frames that constitute the target short video are determined according to the shot planning contained in the director's script. The image frames of each of the aforementioned shots are stitched together in the stitching order, and a transition frame is added between the last frame and the first frame of two consecutive shots with scene changes after stitching, according to the transition method.
3. The method according to claim 1, wherein, The step of editing the image frames of the target short video based on the storyboard design information contained in the director's script further includes: Based on the character entrance plan contained in the director's script, a subtitle bar describing the newly appearing character is generated for the first image frame in the target short video that presents the character.
4. The method according to claim 3, wherein, The generated subtitle bar has a style that matches the style of the first image frame presenting the new character and / or the style of the new character, and the subtitle bar is placed in a position in the first image frame presenting the new character that does not overlap with any subject.
5. The method according to claim 3, wherein, The step of editing the image frames of the target short video based on the storyboard design information contained in the director's script further includes: Based on the emotional change information of the characters in the director's script, the plot rhythm information is determined; Generate plot rhythm tags for the target short video based on the plot rhythm information; A rhythm index is calculated based on the video duration within different plot rhythm tags, and the rhythm of the target short video is adjusted to the desired balanced rhythm based on the rhythm index.
6. The method according to any one of claims 2-5, further comprising: Identify redundant image frames and / or abnormal frames in the target short video; Redundant image frames are removed from the identified frames. The identified abnormal frames are repaired according to the abnormality type using the corresponding repair method; wherein, the abnormality type includes: still frame, blurred frame, and dialogue misalignment frame.
7. The method according to claim 1, wherein, The step of adding matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script includes: Determine the semantics of the plot keywords extracted from the director's script; Determine existing background sounds and / or existing ambient sounds from a preset sound effects library that match the semantic meaning; In response to the absence of background sound and / or ambient sound matching the semantics in the sound effects library, a new background sound and / or new ambient sound matching the semantics is generated using a preset audio generation model; Add the audio that matches the semantics to the video frame corresponding to the plot keyword.
8. The method according to claim 7, further comprising: Determine the mixing balance strategy for dialogue audio, background sound, and ambient sound in the target short video; The dialogue audio, background sound, and ambient sound are mixed according to the mixing balance strategy.
9. The method according to claim 1, wherein, The step of adding matching special effects to the target short video based on the visual language tags and scene tags contained in the director's script includes: Generate text animations, character effects, background effects, and filter effects based on the visual language tags; The scene change time is determined based on the scene label, and a highlight visual effect is generated for the scene change time.
10. The method of claim 9, further comprising: The plot rhythm is determined based on the scene tags, and the style of each special effect within each plot rhythm is controlled to match the corresponding plot rhythm.
11. The method according to claim 1, wherein, The step of editing image frames of the target short video based on the storyboard design information contained in the director's script includes: Using a pre-set video editing AI agent, the image frames of the target short video are edited based on the storyboard design information contained in the director's script; The step of adding matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script includes: Using a pre-defined audio-adding agent, based on the semantic understanding results of the plot keywords contained in the director's script, matching audio is added to the target short video; The step of adding matching special effects to the target short video based on the visual language and scene tags contained in the director's script includes: Using preset special effects, an intelligent agent adds matching special effects to the target short video based on the visual language and scene tags contained in the director's script.
12. The method according to claim 1, further comprising: A multi-dimensional quality assessment is performed on the edited target short video, and modification suggestions are generated when the multi-dimensional quality assessment fails. Based on the proposed modifications, at least one of the video images, audio, and special effects shall be adjusted until the adjusted target short video passes the multidimensional quality assessment.
13. The method according to claim 12, wherein, The process of performing a multi-dimensional quality assessment on the edited target short video, generating modification suggestions when the multi-dimensional quality assessment fails, and adjusting at least one of the video images, audio, and special effects according to the modification suggestions until the adjusted target short video passes the multi-dimensional quality assessment includes: A preset editing evaluation agent is used to perform a multi-dimensional quality assessment on the edited target short video. If the multi-dimensional quality assessment fails, modification suggestions are generated. Based on the modification suggestions, at least one of the video images, audio, and special effects is adjusted until the adjusted target short video passes the multi-dimensional quality assessment.
14. The method according to claim 12 or 13, wherein, The quality dimensions assessed in the multidimensional quality assessment include at least two of the following: The evaluation criteria for the continuity and consistency of images between two consecutive shots, the degree of audio-visual synchronization, sound field balance, composition score, color distribution score, and rhythmic integrity.
15. A short video editing device, comprising: The director script acquisition unit is configured to acquire the director script of the target short video to be edited; The video editing unit is configured to edit the image frames of the target short video based on the storyboard design information contained in the director's script; An audio adding unit is configured to add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script; The special effects addition unit is configured to add matching special effects to the target short video based on the visual language and scene tags contained in the director's script; The timeline alignment unit is configured to align the edited image frames, the added audio stream, and the effects along the timeline to obtain the edited target short video.
16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the short video editing method according to any one of claims 1-14.
17. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the short video editing method according to any one of claims 1-14.
18. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the short video editing method according to any one of claims 1-14.
Citation Information
Patent Citations
Dynamic weather particle effect processing method and device, equipment and storage medium
CN116258802A
Video post-editing and video synthesis optimization method
CN116847123A
Short video intelligent creation system and method based on AI image
CN119763017A
Interactive video generation method capable of intelligent interaction based on multi-agent cooperation and related device
CN120281978A
AI intelligent short video generation method and system based on multi-agent collaboration
CN120547420A