Short video editing system capable of automatically aligning multi-track time axis

Through multi-track alignment analysis unit and deep learning technology, the precise synchronization and subtitle matching of audio and video in short video editing systems are achieved, solving the problem of inefficiency in traditional methods and improving the accuracy and logic of multi-track editing.

CN120434484APending Publication Date: 2025-08-05南京地平线网络科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510749054.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The multi-track editing function in traditional short video editing is inefficient, making it difficult to accurately synchronize audio and video, and the accuracy of subtitles and audio matching and multi-track conflict handling is poor.

Method used

The multi-track alignment analysis unit is used to accurately segment video and audio tracks, combining deep learning and natural language processing technology to identify matching abnormalities between subtitles and audio, and formulate detailed processing strategies through the alignment conflict analysis unit to resolve conflicts in multi-track editing.

Benefits of technology

It improves the efficiency and accuracy of track alignment, ensures accurate matching of subtitles and audio content, improves the accuracy of video information communication and editing logic, and reduces the cumbersomeness of manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434484A_ABST
    Figure CN120434484A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-track time axis automatic alignment short video editing system, relates to the technical field of short video editing, and solves the technical problems that a traditional method is poor in accuracy, and audio and video synchronization is difficult to realize accurately. According to the invention, the conformity between subtitles and audio semantics is analyzed by using a natural language processing technology through the comprehensive sorting and editing unit, the matching error between the subtitles and the audio can be quickly identified, and abnormal subtitles are intelligently screened and pertinently corrected through character identification of the audio and comparative analysis of the subtitles, so that the accuracy of the subtitles is improved. According to the method, subtitles and audio contents are accurately matched, the accuracy of video information transmission is improved, an alignment conflict analysis unit formulates detailed and scientific processing strategies for different types of conflicts such as timeline dislocation, space overlapping conflicts and logic contradictions, priority rules of multi-track alignment are defined, and the accuracy of video information transmission is improved. The conflict problem in the multi-track editing process can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of short video editing, and in particular to a short video editing system with automatic alignment of multi-track timelines. Background Art

[0002] With the booming development of the short video industry, users have put forward higher requirements on the efficiency and quality of video editing.

[0003] According to the patent application with publication number CN112995756A, a short video generation method and device, and a short video generation system are disclosed. The method includes: segmenting multimedia video content to obtain multiple video clips; performing value evaluation on the multiple video clips to obtain an evaluation value of each of the multiple video clips; selecting some video clips from the multiple video clips whose evaluation values are higher than a predetermined threshold according to the evaluation value of each video clip; and generating short videos based on some video clips.

[0004] However, despite the widespread adoption of multi-track editing in existing short video editing technologies, several shortcomings remain. For one thing, track alignment often relies on manual operation, which is inefficient when dealing with complex scenarios like multi-camera shooting and audio / video separation, and is prone to timeline misalignment. Furthermore, traditional methods lack accuracy in lip syncing and audio / video matching, making it difficult to accurately synchronize audio and video. Furthermore, there are challenges with matching subtitles with audio and resolving conflicts between multiple tracks. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a short video editing system with automatic alignment of multi-track timelines, which solves the problem that traditional methods are inaccurate and difficult to accurately achieve synchronization of audio and video.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a short video editing system with automatic alignment of multi-track timelines, comprising: The multi-track alignment analysis unit is used to obtain the multi-track information transmitted by the short video information acquisition unit and perform alignment analysis. The video track is divided into segmented video tracks based on the video screen switching points, and the audio track is divided into segmented audio tracks based on the duration of the video screen. At the same time, the alignment analysis signal or matching information is generated by matching according to the label order. Processing the generated alignment analysis signal, determining the best alignment point to generate matching information, and transmitting it to the comprehensive arrangement and editing unit; The comprehensive editing unit is used to analyze the obtained matching information to determine whether there are any matching anomalies between the audio track and the subtitle track. If so, the audio track is recognized to generate a text recognition result, which is then corrected again with the subtitle track recognition to generate a processed video, which is then transmitted to the alignment conflict analysis unit. The alignment conflict analysis unit is used to perform conflict analysis on the multi-track information in the processed video, and determine the corresponding conflict type and conflict cause. At the same time, it processes the conflict cause separately to generate decision processing information, and finally transmits it to the editing information output unit.

[0007] As a further solution of the present invention, it also includes a short video information collection unit and an editing information output unit; A short video information acquisition unit, used to transmit the acquired multi-track information to a multi-track alignment analysis unit; The editing information output unit is used to display the obtained decision processing information to the corresponding management personnel.

[0008] As a further solution of the present invention, the multi-track alignment analysis unit performs alignment analysis on the multi-track information in the following specific manner: Get the video track and audio track, then split them using the switching points of the video screen as the segmentation points to obtain segmented video tracks, labeled i, where i=1, 2, ..., j, where j represents the number of segmented video tracks. At the same time, obtain the track video duration Ti corresponding to the segmented video track i, and segment the audio track based on the track video duration Ti to obtain segmented audio tracks, labeled n, where n=1, 2, ..., m, where m represents the number of segmented audio tracks; Get segmented video track i, recognize the lip movements of the characters in its video, and obtain the recognition results. Then match the recognition results of segmented video track i with the audio results in segmented audio track n to determine whether the two with the same label correspond to each other.

[0009] As a further solution of the present invention, the multi-track alignment analysis unit determines whether two tracks with the same label correspond to each other in the following specific manner: The dynamic feature vectors of the characters' mouth shapes in the segmented video are extracted through the deep learning model, and the acoustic feature vectors of the segmented audio are extracted using the speech processing tool. The extracted video and audio features are converted into comparable values (cosine similarity) and the formula is used. The confidence level is calculated as is the weight coefficient, and the cosine similarity range is [-1, 1]; The obtained matching confidence is compared with the threshold, and the specific value of the threshold is set by the operator. If the matching confidence is greater than the threshold, it means the match is successful, and they are aligned to generate matching information. Conversely, if the matching confidence is less than the threshold, it means the match fails, an alignment analysis signal is generated, and secondary analysis processing is performed.

[0010] As a further solution of the present invention, the specific manner in which the multi-track alignment analysis unit performs secondary analysis processing is: Obtain the segmented video and audio tracks corresponding to the alignment analysis signal, search for the best alignment point in a small range, make slight time adjustments to the audio or video, align the two, generate matching information, and transmit the generated matching information to the comprehensive arrangement and editing unit.

[0011] As a further solution of the present invention, the comprehensive arrangement and editing unit analyzes the matching information in the following specific manner: Get the subtitle track and audio track, compare them to see if they match, if the match is wrong, generate subtitle recognition signal and analyze it; if the match is correct, directly edit and generate the processed video, at the same time, perform text recognition on the audio track, compare the result with the subtitle track, find out the abnormal subtitles, if the text recognition result is correct, modify the abnormal subtitles based on it and generate the video; if it is wrong, generate the video based on the original subtitle track and output the matching information.

[0012] As a further solution of the present invention, the specific manner in which the alignment conflict analysis unit performs conflict analysis on the multi-track information in the processed video is: When multiple tracks trigger alignment instructions at the same time, priority rules need to be set to determine whether there is a conflict. If so, a conflict analysis signal is generated. Otherwise, a normal output signal is generated and transmitted to the editing information output unit at the same time. The generated conflict analysis signals are analyzed to determine the types of conflicts and the corresponding causes of conflicts, and are processed according to different conflict causes. The conflict causes include timeline misalignment, spatial overlap conflicts and logical contradictions.

[0013] As a further solution of the present invention, the alignment conflict analysis unit processes the conflict according to different conflict causes in the following specific manner: Analyze the timeline misalignment as the cause of the conflict, lock the core track, split the secondary track material, drag the secondary track material frame by frame to align the keyframes, and finally synchronize the entire system. Analyze the spatial overlap conflict as the cause of the conflict, establish the track hierarchy, adjust the blending mode and transparency of the overlapping elements, and set displacement keyframes for the dynamic elements; Analyze the logical contradictions caused by the conflict, separate the audio and video tracks, locate the conflict points, adjust the audio rhythm or replace the clips to match the mood of the picture, and add transitions or sound effects at the conflict points; After the processing is completed, decision processing information is generated and transmitted to the editing information output unit.

[0014] The present invention provides a short video editing system with automatic alignment of multiple track timelines. Compared with the existing technology, it has the following advantages: The present invention uses a multi-track alignment analysis unit, combined with three algorithms: inter-frame difference detection, color histogram change, and deep learning scene classification, to determine the video transition points, thereby achieving accurate segmentation of video and audio tracks. Based on the deep learning "audio-lip shape mapping" method, it comprehensively uses multiple models and technologies for lip shape recognition, and judges the degree of audio and video matching by calculating the matching confidence, which greatly improves the efficiency and accuracy of track alignment and reduces the tediousness of manual operation.

[0015] The present invention uses natural language processing technology to analyze the semantic consistency of subtitles and audio through a comprehensive editing unit, which can quickly identify matching errors between subtitles and audio. It also intelligently filters out abnormal subtitles and performs targeted corrections through text recognition of audio and comparative analysis with subtitles, ensuring that subtitles are accurately matched with audio content and improving the accuracy of video information transmission.

[0016] The present invention uses the alignment conflict analysis unit to formulate detailed and scientific processing strategies for different types of conflicts such as timeline misalignment, spatial overlap conflict and logical contradiction, and clarifies the priority rules for multi-track alignment. It can solve the conflict problems in the multi-track editing process and ensure the logic of video editing and the smoothness of visual effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a block diagram of the system principle of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] Example 1 See also Figure 1 The present application provides a short video editing system with automatic alignment of multi-track timelines, including a short video information acquisition unit, a multi-track alignment analysis unit, a comprehensive editing unit, an alignment conflict analysis unit and an editing information output unit, and combines Figure 1It can be known that the functional units are electrically connected in a unidirectional manner.

[0020] The short video information acquisition unit is used to acquire the target video and the multi-track information of the target video at the same time. The multi-track information includes video track, audio track and subtitle track, and transmits it to the multi-track alignment analysis unit.

[0021] A multi-track alignment analysis unit is used to perform track alignment analysis on the target video according to the obtained multi-track information, obtain the video track and audio track, and the audio track here is represented by the speaking audio of the characters in the video, and then use the transition point of the video track as the segmentation point, and the transition point is represented by the switching point of the video screen, and combine three transition detection algorithms to determine the switching point, inter-frame difference detection (detecting mutations), color histogram changes (detecting gradual changes) and deep learning scene classification (identifying scene changes), and segment it to obtain segmented video tracks labeled i, and i=1, 2, ..., j, where j represents the number of segmented video tracks, and at the same time obtain the track video duration Ti corresponding to the segmented video track i, and segment the audio track based on the track video duration Ti to obtain segmented audio tracks, specifically here the starting point of the video track is used as the starting point of the audio track for segmentation, and the label is n, and n=1, 2, ..., m, where m represents the number of segmented audio tracks; Then, segmented video track i is matched with segmented audio track n in the order of numbers, and the specific matching method is as follows: Obtain segmented video track i and simultaneously identify the lip shapes of the characters in the video. The identification here is based on the "audio-lip shape mapping" method in the deep learning method. Combined with FaceAlignment, FacialLandmarkDetection and CNN model, image enhancement technology is used to highlight lip details, adjust the lip ROI in real time, adapt to head movement, capture the 3D deformation characteristics of the lips, and obtain recognition results. Then, the recognition results of segmented video track i are matched with the audio results in segmented audio track n to determine whether the two with the same labels correspond to each other. The matching judgment method here is analyzed by calculating the matching confidence of the two. The specific analysis method is to extract the dynamic feature vector of the character's lip shape in the segmented video through a deep learning model (such as 3D CNN, OpenFace) (lip shape opening and closing degree (such as lip width and height), facial expression key point coordinates (such as mouth corners, jaw movement trajectory), lip shape duration sequence (such as the number of frames of each lip shape state)), and use speech processing tools (such as Librosa, PyTorch Speech) extracts the acoustic feature vector of the segmented audio, converts the extracted video and audio features into comparable values (cosine similarity) and uses the formula The confidence level is calculated as is the weight coefficient, the cosine similarity range is [-1, 1], here it is normalized to [0, 1]. The smaller the DTW distance, the higher the matching degree, which is converted to a score in the [0, 1] interval through normalization; The obtained matching confidence is compared with a threshold value, and the specific value of the threshold value is set by the operator. If the matching confidence is greater than the threshold value (>), it indicates a successful match, and the two are aligned to generate matching information. Conversely, if the matching confidence is less than the threshold value (≤), it indicates a matching failure, and a secondary analysis is performed. At the same time, an alignment analysis signal is generated. This is repeated for all segmented video tracks i and segmented audio tracks n. The generated alignment analysis signal is then processed to obtain the segmented video and audio tracks corresponding to the alignment analysis signal. The optimal alignment point is searched within a small range (±0.5 seconds), and a slight time adjustment (≤5%) is made to the audio or video. The two are then aligned to generate matching information, which is then transmitted to the comprehensive editing unit.

[0022] A comprehensive editing unit is used to analyze the obtained matching information, obtain the subtitle track, and identify the subtitle track based on the audio track in the matching information. Natural language processing technology is used to analyze whether the subtitle text and the audio semantics are consistent, and to determine whether there is a mismatch between the two. Here, a mismatch indicates that the subtitle track and the character audio do not match. If so, a subtitle recognition signal is generated. Otherwise, the two are matched and comprehensively edited to generate a processed video, and the generated subtitle recognition signal is analyzed. Perform text recognition on the audio track in the matching information to generate a text recognition result. At the same time, compare the text recognition result with the subtitle track, screen out subtitles with matching anomalies and record them as abnormal subtitles, then judge the correctness of the text recognition result. If the text recognition result is correct, modify and match the abnormal subtitles based on the text recognition result to generate a processed video. Conversely, if the text recognition result is abnormal, match the subtitle track based on the standard to generate a processed video, and transmit the generated matching information to the editing information output unit.

[0023] The editing information output unit is used to display the acquired processing video to the corresponding management personnel.

[0024] Example 2 As the second embodiment of the present invention, it is implemented on the basis of the first embodiment, and differs from the first embodiment in the following aspects: The comprehensive editing unit transmits the generated processed video to the alignment conflict analysis unit for alignment conflict analysis. This unit performs conflict analysis on the multi-track information in the processed video. Specifically, when multiple tracks trigger alignment instructions simultaneously (e.g., the video track needs to be aligned according to picture motion, and the audio track needs to be aligned according to drum beats), priority rules need to be set. For example, in movie editing, the dialogue track has a higher priority than the background music track, and the vocals and lip movements need to be aligned first. A conflict is determined. If so, a conflict analysis signal is generated. Otherwise, a normal output signal is generated and transmitted to the editing information output unit. Then, the generated conflict analysis signal is analyzed to determine the conflict type and the corresponding conflict cause. Different conflict causes are handled according to the conflict. The conflict causes include timeline misalignment, spatial overlap conflict, and logical contradiction. The specific handling methods are as follows: If the cause of the conflict is a timeline misalignment, analyze the process. First, identify the core track and lock it. Then, use editing tools to split the secondary track material into conflicting and non-conflicting segments. Delete or move the conflicting segments as needed. Next, zoom in on the timeline and drag the secondary track material frame by frame to align keyframes (such as lip sync, action peaks, and sound effects and drum beats) with the main track. Finally, synchronize the entire track. Analyze the cause of the conflict as spatial overlap, establish a track hierarchy, place the main video track at the bottom of the timeline panel, and add layers in sequence according to the track hierarchy. Then adjust the layer blending mode and transparency (use a blending mode (such as Screen or Overlay) for overlapping elements or reduce transparency (such as setting the transparency of subtitles to 80%) to reduce the sense of obstruction). Finally, set displacement keyframes for dynamic elements (such as scrolling subtitles and floating stickers); If the cause of the conflict is a logical contradiction, analyze it, separate the audio and video tracks, and decouple the video track from the audio track (shortcut key Ctrl+Shift+G). Listen to the audio or preview the video separately to locate the logical contradiction. Then, match the mood and rhythm, adjust the rhythm of the background music (such as speeding up / slowing down) or replace the music clip to synchronize it with the mood of the picture (such as tension, joy). Finally, add transitions (such as fades) or sound effects (such as "ding") at the conflict points (such as scene changes and sudden changes in audio and video) to weaken the sense of logical discontinuity. Decision processing information is generated based on the above-mentioned processing and is simultaneously transmitted to the editing information output unit.

[0025] The editing information output unit is used to display the acquired decision processing information to the corresponding management personnel.

[0026] Example 3 As the third embodiment of the present invention, the focus is on combining the implementation processes of the first and second embodiments.

[0027] Some of the data in the above formulas are calculated based on their numerical values and are not substituted into parameter units for calculation. At the same time, the contents not described in detail in this specification belong to the existing technology known to those skilled in the art.

[0028] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A short video editing system with automatic alignment of multi-track timelines, characterized by: include: The multi-track alignment analysis unit is used to obtain the multi-track information transmitted by the short video information acquisition unit and perform alignment analysis. The video track is divided into segmented video tracks based on the video screen switching points, and the audio track is divided into segmented audio tracks based on the duration of the video screen. At the same time, the alignment analysis signal or matching information is generated by matching according to the label order. Processing the generated alignment analysis signal, determining the best alignment point to generate matching information, and transmitting it to the comprehensive arrangement and editing unit; The comprehensive editing unit is used to analyze the obtained matching information to determine whether there are any matching anomalies between the audio track and the subtitle track. If so, the audio track is recognized to generate a text recognition result, which is then corrected again with the subtitle track recognition to generate a processed video, which is then transmitted to the alignment conflict analysis unit. The alignment conflict analysis unit is used to perform conflict analysis on the multi-track information in the processed video, and determine the corresponding conflict type and conflict cause. At the same time, it processes the conflict cause separately to generate decision processing information, and finally transmits it to the editing information output unit.

2. The short video editing system with automatic multi-track timeline alignment according to claim 1, characterized in that: It also includes a short video information collection unit and an editing information output unit; A short video information acquisition unit, used to transmit the acquired multi-track information to a multi-track alignment analysis unit; The editing information output unit is used to display the obtained decision processing information to the corresponding management personnel.

3. The short video editing system with automatic multi-track timeline alignment according to claim 1, characterized in that: The specific method of the multi-track alignment analysis unit to perform alignment analysis on the multi-track information is as follows: Get the video track and audio track, then split them using the switching points of the video screen as the segmentation points to obtain segmented video tracks, labeled i, where i=1, 2, ..., j, where j represents the number of segmented video tracks. At the same time, obtain the track video duration Ti corresponding to the segmented video track i, and segment the audio track based on the track video duration Ti to obtain segmented audio tracks, labeled n, where n=1, 2, ..., m, where m represents the number of segmented audio tracks; Get segmented video track i, recognize the lip movements of the characters in its video, and obtain the recognition results. Then match the recognition results of segmented video track i with the audio results in segmented audio track n to determine whether the two with the same label correspond to each other.

4. The short video editing system with automatic multi-track timeline alignment according to claim 3, characterized in that: The specific method of the multi-track alignment analysis unit to determine whether two items with the same label correspond to each other is as follows: The dynamic feature vectors of the characters' mouth shapes in the segmented video are extracted through the deep learning model, and the acoustic feature vectors of the segmented audio are extracted using the speech processing tool. The extracted video and audio features are converted into comparable values (cosine similarity) and the formula is used. The confidence level is calculated as is the weight coefficient, and the cosine similarity range is [-1, 1]; The obtained matching confidence is compared with the threshold, and the specific value of the threshold is set by the operator. If the matching confidence is greater than the threshold, it means the match is successful, and they are aligned to generate matching information. Conversely, if the matching confidence is less than the threshold, it means the match fails, an alignment analysis signal is generated, and secondary analysis processing is performed.

5. The short video editing system with automatic multi-track timeline alignment according to claim 4, characterized in that: The specific method of the multi-track alignment analysis unit performing secondary analysis is as follows: Obtain the segmented video and audio tracks corresponding to the alignment analysis signal, search for the best alignment point in a small range, make slight time adjustments to the audio or video, align the two, generate matching information, and transmit the generated matching information to the comprehensive arrangement and editing unit.

6. The short video editing system with automatic multi-track timeline alignment according to claim 1, characterized in that: The specific method for the comprehensive arrangement and editing unit to analyze the matching information is as follows: Obtain the subtitle track and audio track, compare them to see if they match, and if not, generate and analyze subtitle recognition signals; if they match correctly, directly edit and generate a processed video. At the same time, perform text recognition on the audio track and compare the results with the subtitle track to identify any abnormal subtitles. If the text recognition results are correct, modify the abnormal subtitles based on them and generate the video. If there is an error, the video will be generated based on the original subtitle track and the matching information will be output.

7. The short video editing system with automatic multi-track timeline alignment according to claim 1, characterized in that: The specific method of the alignment conflict analysis unit performing conflict analysis on the multi-track information in the processed video is as follows: When multiple tracks trigger alignment instructions at the same time, priority rules need to be set to determine whether there is a conflict. If so, a conflict analysis signal is generated. Otherwise, a normal output signal is generated and transmitted to the editing information output unit at the same time. The generated conflict analysis signals are analyzed to determine the types of conflicts and the corresponding causes of conflicts, and are processed according to different conflict causes. The conflict causes include timeline misalignment, spatial overlap conflicts and logical contradictions.

8. The short video editing system with automatic multi-track timeline alignment according to claim 7, characterized in that: The specific way in which the alignment conflict analysis unit processes the conflict according to different causes is as follows: Analyze the timeline misalignment as the cause of the conflict, lock the core track, split the secondary track material, drag the secondary track material frame by frame to align the keyframes, and finally synchronize the entire system. Analyze the spatial overlap conflict as the cause of the conflict, establish the track hierarchy, adjust the blending mode and transparency of the overlapping elements, and set displacement keyframes for the dynamic elements; Analyze the logical contradictions caused by the conflict, separate the audio and video tracks, locate the conflict points, adjust the audio rhythm or replace the clips to match the mood of the picture, and add transitions or sound effects at the conflict points; After the processing is completed, decision processing information is generated and transmitted to the editing information output unit.

Citation Information

Patent Citations

  • Multi-channel video signal alignment method and device and electronic equipment

    CN115460446A

  • Short video editing method and system based on artificial intelligence

    CN119031197A

  • Automatic synchronization between content video and subtitle using artificial intelligence

    KR102555698B1

  • System and method of automatically aligning video scenes with an audio track

    US7512886B1

  • Recording medium recorded with multi-track media file, method for editing multi-track media file, and apparatus for editing multi-track media file

    WO2014181969A1