Automatic video editing processing method and device based on sound and optical flow, and terminal

Through the automated video editing method based on sound and optical flow, the problem of insufficient audio-visual fragmentation and dynamic adaptability is solved, and the joint analysis and processing of audio and video information is realized, high-quality short videos are generated, and the mobile terminal or low-configuration equipment is adapted.

CN120378713APending Publication Date: 2025-07-25SHENZHEN KUKAI SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510576456.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing video automation editing technology has problems such as audio-visual fragmentation, insufficient dynamic adaptability and imbalance in efficiency and effect, and it is difficult to realize the joint analysis and processing of audio and video information, especially in the rhythm extraction of complex motion scenarios and non-periodic sounds, and it is difficult to adapt to mobile terminals or low-configuration devices.

Method used

An automated video editing method based on sound and optical flow is adopted, and through audio and video separation, music interval and rhythm point recognition, lens sharding, frame extraction analysis, background sound energy calculation and other steps, combined with deep learning models, audio and video synchronization and high-ignition lens selection are generated to generate high-quality short videos.

Benefits of technology

The joint modeling of dynamic correlation between sound and picture is realized, dynamic adaptability and editing effect are improved, and high-quality short videos can be generated efficiently on mobile terminals or low-configuration devices, attracting users' attention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378713A_ABST
    Figure CN120378713A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic video editing processing method, device and terminal based on sound and optical flow, and belongs to the technical field of video processing, and the method comprises the steps: obtaining a to-be-edited video, and carrying out the audio-video separation; performing music interval identification and rhythm point identification on the input specified music; performing shot fragmentation on the video medium; performing frame extraction analysis on the video medium, and calculating shot motion data information and subtitle information; separating background sound from the audio medium and calculating an energy value; the end time of the mixed and clipped video is calculated based on the music interval recognition result, a specified shot is selected as the head of the clip, a high-combustion shot is sorted and selected based on shot motion data information to serve as a clip in the clip, an end word is selected based on subtitle information to find out a corresponding shot to serve as the tail of the clip, video assembly is conducted, and a video logo is added to generate an editing result. According to the invention, automatic processing of video editing is realized, and the video editing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly relates to an automated video editing processing method, device, intelligent terminal, and storage medium based on sound and optical flow. Background Art

[0002] With the rise and popularity of short video platforms, video editing technology plays an increasingly important role in the field of content creation. As a technical means to improve the efficiency of video production, automated video editing has become a hot topic in current research. Currently, automated video editing technology mainly focuses on aspects such as video segmentation, content recognition, and audio-video synchronization.

[0003] However, the existing video automated editing technology still has the following problems: First, the existing technology often adopts an independent processing method when dealing with audio and video signals, lacking joint modeling of the dynamic association between sound and picture, resulting in the problem of audio-visual disconnection. Especially when it is necessary to perform video beat editing according to the music rhythm points, the existing technology is difficult to accurately identify the music intervals and rhythm points and cannot achieve high-quality audio-video synchronization.

[0004] Second, the existing technology has deficiencies in dynamic adaptability. Traditional optical flow algorithms are prone to calculation deviations in complex motion scenes, and the audio analysis module has weak rhythm extraction ability for non-periodic sounds, making it difficult to accurately capture high-energy shots and music high-energy intervals in the video, affecting the editing effect.

[0005] Third, it is difficult for the existing technology to achieve a balance between efficiency and effect. Although the multi-modal scheme based on deep learning has good effects, it usually relies on high-computing-power GPUs for real-time inference, making it difficult to adapt to mobile devices or low-configuration devices, restricting the application scenarios of the technology. In addition, there is still room for improvement in the automation level of the existing technology, especially in aspects such as lens segmentation, motion data calculation, and subtitle information extraction, lacking systematic solutions.

[0006] Therefore, there is an urgent need for an automated video editing processing method based on sound and optical flow that can effectively solve the above problems, realize the joint analysis and processing of audio-video information, and improve the automation level and editing effect of video editing. Summary of the Invention

[0007] In order to solve the technical problems such as the audio-visual disconnection problem, insufficient dynamic adaptability, and imbalance between efficiency and effect existing in the existing video automated editing technology, and to achieve the joint modeling of dynamic association between sound and picture, improve dynamic adaptability, and balance efficiency and effect, the present invention provides an automated video editing processing method, device, intelligent terminal, and storage medium based on sound and optical flow. The present invention provides an automated editing method that can deeply integrate audio-visual features, is lightweight, and has strong dynamic adaptability. The automated video editing based on sound and optical flow of the present invention generates high-quality short videos, attracts users' attention, and improves the efficiency and quality of video generation.

[0008] The technical solutions adopted by the present invention to solve the problems are as follows: Provide an automated video editing processing method based on sound and optical flow, including: obtaining the video to be edited, separating the audio and video of the video to be edited to obtain an audio medium and a video medium respectively; identifying the music interval of the input specified music, then identifying the rhythm points, and finding the positions where the music makes a beat; segmenting the video medium into shots and extracting all the shots of the video; performing frame extraction analysis on the video medium, extracting each frame of the video and calculating the lens movement data information and subtitle information; separating the background sound from the audio medium and calculating the energy value of the background sound; performing mixed video assembly: based on the music interval identification result, calculating the end time of the mixed video; selecting a specified shot from the extracted shots as the intro; sorting based on the calculated lens movement data information and selecting high-energy shots located in the high-energy interval of the sound as the middle segments; selecting the ending words based on the subtitle information and finding the corresponding shots as the ending; performing video assembly and adding a video logo to generate the video editing result.

[0009] Preferably, the music interval analysis includes analyzing the prelude, verse, pre-chorus, chorus, interlude, bridge, and coda of the audio medium.

[0010] Further, the step of segmenting the video medium into shots and extracting all the shots of the video includes: segmenting the video medium separated from the video to be edited into shots and extracting all the shots of the video medium, where each section between two scene switches is regarded as a shot.

[0011] Further, the step of performing frame extraction analysis on the video medium, extracting each frame of the video and calculating the lens movement data information and subtitle information includes: performing frame extraction analysis on the video medium, extracting each frame of the video and calculating the lens movement data information and subtitle information, where the lens movement data information is the active movement data shown during the lens switching process, and the subtitle information is used to provide data support for the large model to select the ending words.

[0012] Further, the step of selecting a specified shot from the extracted shots as the opening title includes: based on all the shots of the extracted video, selecting the shots with the subject tag of scenery as the opening title; agreeing on the total duration of the opening title, restricting the time of each shot not to exceed the specified time, and realizing that the music beat point is the shot transition point, and eliminating the shots with subtitles.

[0013] Further, the step of sorting based on the calculated shot motion data information and selecting the high-energy shots located in the high-energy interval of the sound as the middle part of the video includes: based on the calculated shot motion data information, sorting the shot motion data information in descending order; according to the sorting of the shot motion data information, screening out the high-energy shots with the sorted motion data information in the front and eliminating the shots with subtitles as the high-energy shots in the middle part of the video.

[0014] Further, the step of selecting the ending word based on the subtitle information and finding the corresponding shot as the ending of the video includes: obtaining the extracted subtitle information; inputting the subtitle information into a large language model for processing to select the ending word based on the large language model; and finding the segment corresponding to the ending word, retrieving and intercepting the corresponding shot as the ending shot of the video; the step of assembling the video and adding the video logo to generate the video editing result includes: according to the calculated ending time of the mixed video, assembling the selected shot as the opening title, the high-energy shots in the middle part of the video, and the corresponding shots as the ending part of the video, and adding audio and video effects and adding the video logo to generate the video editing result.

[0015] An automated video editing processing device based on sound and optical flow, wherein the device includes: An audio-video separation module, configured to obtain the video to be edited, perform audio-video separation on the video to be edited, and respectively obtain an audio medium and a video medium; A music interval recognition module, configured to perform music interval recognition on the input specified music, then perform rhythm point recognition, and find the positions where the music beats; A shot segmentation module, configured to segment the video medium into shots and extract all the shots of the video; A frame extraction analysis module, configured to perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the shot motion data information and subtitle information; A background sound energy value calculation module, configured to separate the background sound from the audio medium and calculate the energy value of the background sound; The mixed video assembly module is used for assembling mixed videos: based on the music interval recognition result, calculate the end time of the mixed video; select a specified shot from the extracted shots as the opening title; sort based on the calculated shot motion data information, and select high-energy shots located in the high-energy interval of the sound as the middle segments; select the ending words based on the subtitle information, and find the corresponding shots as the ending; perform video assembly and add a video logo to generate the video editing result.

[0016] An intelligent terminal, which includes a memory and one or more programs. One or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing any of the described methods.

[0017] A computer-readable storage medium, wherein when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute any of the described methods.

[0018] The beneficial effects of the present invention are as follows: By performing music interval recognition and rhythm point recognition on the specified input music, and combining the shot segmentation and motion data information analysis of the video medium, the present invention realizes the joint modeling of the dynamic association between sound and picture, and solves the problem of audio-visual disconnection in traditional methods; By performing frame extraction analysis on the video medium to calculate the shot motion data information and calculating the background sound energy value of the audio medium, the dynamic adaptability is improved, and the optical flow calculation in complex motion scenes and the rhythm extraction of non-periodic sounds can be effectively processed; The method of the present invention does not need to rely on high-computing-power GPU real-time inference, and balances efficiency and effect through preprocessing and intelligent assembly, and can be adapted to mobile devices or low-configuration devices; Through precise audio-visual synchronization and dynamic editing, high-quality short videos can be generated, effectively attracting users' attention. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of the automated video editing processing method based on sound and optical flow provided in Embodiment 1 of the present invention.

[0021] Figure 2 It is a schematic diagram of the high-energy video interval structure of the automated video editing processing method based on sound and optical flow provided in Embodiment 2 of the present invention.

[0022] Figure 3 It is a schematic flowchart of the automated video editing processing method based on sound and optical flow provided in Embodiment 4 of the present invention.

[0023] Figure 4 It is a principle block diagram of the embodiment of the automated video editing processing device based on sound and optical flow provided by the present invention.

[0024] Figure 5 It is a principle block diagram of the internal structure of the intelligent terminal provided in the embodiment of the present invention. Specific embodiments

[0025] To make the objectives, technical solutions and advantages of the present invention clearer and more explicit, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0026] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If this specific posture changes, the directional indications will also change accordingly.

[0027] With the explosive growth of the short video industry, content creators have an increasingly urgent need for efficient and intelligent automated editing tools. Existing video automated editing technologies mainly rely on key frame detection, scene segmentation, or simple synchronization technologies based on audio beats. For example, some tools extract the peak of the audio spectrum to match the video switching point, or use object detection algorithms to segment the scene for editing. However, the existing technologies have the following significant defects: Audio-visual disconnection: Traditional methods mostly process audio and video signals independently, lacking joint modeling of the dynamic association between sound and picture. For example, cutting the picture only based on the beat may result in discontinuous actions, while scene segmentation relying only on the optical flow method ignores the impact of sound emotion on the editing rhythm.

[0028] Insufficient dynamic adaptability: Existing optical flow algorithms are prone to calculation deviations in complex motion scenes (such as rapid shot switching and multi-object motion), resulting in inaccurate selection of editing points; the audio analysis module has weak rhythm extraction ability for non-periodic sounds (such as ambient sound and dialogue).

[0029] Imbalance between efficiency and effect: Although multi-modal solutions based on deep learning can improve the effect, they rely on high-computing-power GPUs for real-time inference, making it difficult to adapt to mobile devices or low-configuration devices, which limits the technical implementation scenarios.

[0030] In this context, the industry urgently needs an automated video editing technology that can deeply integrate audio-visual features, is lightweight, and has strong dynamic adaptability. The automated video editing technology based on sound and optical flow of the present invention can generate high-quality short videos with high generation efficiency.

[0031] Embodiment 1 As Figure 1 shown, an automated video editing and processing method based on sound and optical flow according to Embodiment 1 of the present invention includes the following steps: Step S100: Obtain the video to be edited, perform audio-visual separation on the video to be edited, and obtain an audio medium and a video medium respectively; Specifically, the present invention can obtain the original video file uploaded by the user that needs to be edited through a video processing system. The video file can be in common video formats such as MP4, AVI, MOV, etc. After the system receives the video file, it calls an audio-visual separation algorithm to decode the video, and extracts the audio data and video data in the video respectively. The audio data is saved as an audio format file such as WAV or MP3 as the audio medium; the video data is saved as a silent video file as the video medium. This separation processing enables the system to independently analyze and process audio and video content, laying a foundation for subsequent music section recognition and video analysis.

[0032] Step S200: Identify the music sections of the input specified music, and then identify the rhythm points to find the positions where the music makes beats; In this step, the user can first select a specified music for video editing, that is, the specified music used as the background music for video editing; select the specified music and input it into the system. The system of the present invention will perform music section recognition on the input specified music. Music section recognition includes analyzing different music structure parts such as the prelude, verse, guide song, chorus, interlude, bridge, coda, etc. of the input specified music. The system identifies different music sections by analyzing audio features such as the spectrum characteristics, energy changes, and pitch changes of the specified music.

[0033] Specifically, when implemented, the system uses a deep learning model to perform segmented analysis on the audio, extracts audio features such as the Mel-frequency cepstral coefficients (MFCC), chroma features, and beat features of the audio of the specified music, and divides the audio into different music sections through a trained music structure recognition model. For example, the system can identify that the prelude part of the music is usually located at 0-30 seconds at the beginning of the music, the verse part may be located at 30-60 seconds, the chorus part may be located at 60-90 seconds, etc.

[0034] After the music section recognition of the specified music is completed, the system further performs rhythm point recognition. Rhythm point recognition is to find the time points suitable for video switching in the music, namely the so-called "beat-matching" positions, by analyzing features such as the energy change, volume peak, and beat information of the audio. These beat-matching positions are usually the positions where the rhythm in the music changes significantly, such as drum beats, accents, and music turning points.

[0035] The system adopts short-time energy analysis and beat detection algorithms to calculate the energy change curve of the audio on the time axis, and identify the energy mutation points and the positions of strong beats. These positions are marked as potential beat-matching positions and sorted according to their importance. For example, the start position of the chorus, the first drum beat in the dense drum area, and the start point of the music climax are all high-quality beat-matching positions.

[0036] The system saves the recognized music section information and beat-matching position information as structured data to provide a music timeline reference for subsequent video editing.

[0037] Step S300: Segment the video medium into shots and extract all the shots of the video; In this step, the system segments the video medium separated from the video to be edited and extracts all the shots of the video medium. Among them, the section between every two scene switches is regarded as a shot.

[0038] Specifically, the system uses a scene change detection algorithm to analyze the video. First, the system uniformly samples the video, extracts key frames, and then calculates the visual difference degree between adjacent key frames. When the difference degree between two adjacent key frames exceeds the preset threshold, the system determines that a scene change has occurred.

[0039] The system uses multiple features to calculate the inter-frame difference, including color histogram difference, edge feature difference, optical flow feature difference, etc. To improve the detection accuracy, the system also adopts an adaptive threshold technology to dynamically adjust the judgment threshold according to the complexity of the video content.

[0040] After detecting the scene change points, the system marks the video segments between every two adjacent change points as an independent shot and assigns a unique identifier to each shot. The system also records basic information such as the start time, end time, and duration of each shot.

[0041] In this way, the original video is segmented into multiple independent shot segments, providing a basic material library for subsequent shot selection and assembly. For example, a 5-minute original video may be segmented into 30 - 50 different shot segments, and each segment represents a continuous scene or action.

[0042] Step S400: Perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the camera movement data information and subtitle information. In this step, perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the camera movement data information and subtitle information. Among them, the camera movement data information is the active movement data shown during the camera switching process, and the subtitle information provides data support for the large model to select the ending words.

[0043] Specifically, the system first evenly extracts frames from the video, usually extracting 5 - 10 frames per second for analysis. For each frame of the image, the system performs two aspects of analysis: On the one hand, the system calculates the camera movement data information. The system uses the optical flow algorithm to calculate the motion vector field between adjacent frames, and analyzes the moving direction and speed of the objects in the picture. The system calculates the amplitude and direction distribution of the optical flow vectors, extracts features such as motion intensity, motion complexity, and motion direction consistency, and quantifies the activity level of the camera movement.

[0044] For example, the system can calculate indicators such as the average motion amplitude of each frame, the variance of the motion vectors, and the main motion direction, and comprehensively evaluate the motion characteristics of the camera. Cameras that move quickly, contain violent actions, or have segments with rapid panning, tilting, or zooming usually receive higher motion data scores.

[0045] On the other hand, the system extracts the subtitle information of each frame. The system uses optical character recognition (OCR) technology to detect and recognize the text content in the video frame. The system first locates the area that may contain text through image processing technology, and then uses the OCR engine to convert the text in the image into editable text information.

[0046] The system not only extracts the text content of the subtitles, but also records the appearance time, duration, position information, etc. of the subtitles. These subtitle information are stored in a structured manner for subsequent content analysis and ending word selection.

[0047] In this way, the system establishes a detailed motion feature file and subtitle content file for each camera, providing data support for subsequent high - energy camera screening and ending word selection.

[0048] Step S500: Separate the background sound from the audio medium and calculate the energy value of the background sound. In this step, the system further processes the previously separated audio medium, separates the background sound, and calculates the energy value of the background sound.

[0049] In specific implementation, the system first processes the audio medium using an audio source separation algorithm to separate the audio into a vocal part and a background music part. The system can adopt a deep learning-based sound source separation model, such as U-Net or Spleeter, etc., to effectively separate the vocals and background music in the mixed audio through spectral analysis.

[0050] After obtaining the background sound, the system calculates the energy distribution of the background sound on the time axis. The system divides the background sound into small time windows (usually 20 - 50 milliseconds), and calculates the short-time energy value for each window. Among them, an example of the short-time energy calculation formula is: E = Σ(x[6]²); where x[6] is the audio sampling point within the window, and E is the energy value of this window.

[0051] The system smooths the calculated energy value sequence to obtain the energy curve of the background sound. The system further analyzes the energy curve to identify the high-energy peak regions, which usually correspond to the climax parts or the parts with strong rhythms of the music, and are marked as "high-energy intervals".

[0052] The system also calculates features such as the average energy value, maximum energy value, and energy change rate of the background sound, which are used to evaluate the overall intensity and dynamic range of the music. Combining these energy features with the previously identified music interval information can more accurately locate the high-energy intervals in the music, providing a basis for subsequent matching of high-energy shots.

[0053] Step S600: Assemble the mixed video: Based on the music interval recognition result, calculate the end time of the mixed video; select a specified shot from the extracted shots as the start of the video; sort based on the calculated shot movement data information, and select high-energy shots located in the high-energy intervals of the sound as the middle segments of the video; select the ending words based on the subtitle information, and find the corresponding shots as the end of the video; assemble the video and add the video logo to generate the video clip result.

[0054] In this step, the system performs the assembly work of the mixed video based on various analysis results obtained in the previous steps.

[0055] First, the present invention calculates the end time of the mixed video based on the music interval recognition result. The system of the present invention selects a complete music structure as the duration of the mixed video, for example, from the start of the music to the end of the first chorus, or from the start of the music to the end of the whole song. The system determines the end time point of the mixed video according to the integrity of the music structure and the expected video length (usually ranging from 15 seconds to 3 minutes).

[0056] Secondly, the system selects a designated shot from the extracted shots as the opening. Based on all the shots of the extracted video, the system selects shots with the subject label of scenery as the opening. The system uses computer vision algorithms to analyze the content of each shot and identify the main objects and scene types in the shot. The system gives priority to shots containing open scenes such as natural scenery and urban landscapes as the opening. Such shots usually have an introductory and foreshadowing nature and are suitable as the opening of the video.

[0057] The system agrees on the total length of the opening, which is usually 10%-15% of the total length of the video. The length of each opening shot is limited to a specified time (usually 2-3 seconds), and the system will use the music card position as the shot switching point to achieve the effect of audio and video synchronization. The system will also remove shots containing subtitles to maintain the purity and visual impact of the opening.

[0058] Next, the system sorts the shots based on the calculated shot motion data information and selects the high-energy shots in the high-energy sound range as the clips in the film. The system sorts all shots from high to low according to the scores of the shot motion data information. The system gives priority to shots with higher motion data scores, which usually contain more dynamic elements and visual impact.

[0059] The system matches high-scoring shots with previously identified high-energy sections of the music, ensuring that the parts with strong music rhythms and high energy are matched with shots with equally strong visual effects. The system also removes shots with subtitles to prevent subtitles from interfering with the visual experience. These high-energy shots are selected as clips in the film and constitute the main part of the video.

[0060] Then, the system selects the ending words based on the subtitle information and finds the corresponding shots as the ending. The system obtains the previously extracted subtitle information and inputs these subtitle texts into the large language model for processing. The large language model selects text fragments from the subtitles that are suitable as the video ending, namely the "ending words", based on semantic understanding and sentiment analysis.

[0061] Ending words usually have the characteristics of summarizing, pointing out the theme, or elevating emotions. The system finds the video clip corresponding to the ending words, retrieves and captures the shots containing the subtitles as the ending shots. If there are multiple shots containing the same or similar ending words, the system will select the shot with higher picture quality and stronger emotional expression.

[0062] Finally, the system assembles the selected shots as the opening, the high-energy shots in the middle, and the corresponding shots as the ending according to the calculated end time of the mixed video. The system splices these shots together in chronological order, ensuring that the shot switching points are aligned with the music card points to achieve the effect of audio and video synchronization.

[0063] The system will also add audio and video effects, such as adding transition effects, adjusting color styles, adding filters, etc., to enhance the overall visual perception of the video. The system will also add a video logo at an appropriate position in the video (usually at the beginning or end) to identify the creator or brand information of the video.

[0064] For example Figure 2 As shown, for the picture (video) part, a scenic shot is selected as the opening title, high-energy shots are mixed and edited in the middle part of the video, and an ending word shot is selected for the ending. For the first part of the audio, the background music (BGM) in the first section corresponding to the opening title and the middle part of the video is selected, and the original sound of the film (video) is reduced in volume; for the ending of the audio, corresponding to the ending part of the video, the original sound of the film is selected at the end of the first chorus of the background music (BGM), and the background music (BGM) is reduced in volume.

[0065] Through the above steps, the system generates the final video editing result, a mixed video based on the original video material, organized according to the music rhythm and with high visual impact.

[0066] Embodiment 2 An automated video editing and processing method based on sound and optical flow provided in this Embodiment 2 includes: Step S11: Obtain the video to be edited, separate the audio and video of the video to be edited, and obtain an audio medium and a video medium respectively; Specifically, the system receives the original video file uploaded by the user, and this video file can be in a variety of common formats. The system calls a professional audio and video separation library, such as FFmpeg, to decode the video file and separate the audio stream and video stream. The audio stream is saved as an independent audio file as the audio medium; the video stream is saved as a silent video file as the video medium. This separation enables the system to process audio and video content separately, providing a basis for subsequent analysis.

[0067] Step S12: Identify the music intervals of the input specified music, then identify the rhythm points, and find the positions where the music makes a beat; In this embodiment, the system conducts a more detailed music interval analysis on the input specified music for music interval identification, including analyzing the complete music structure such as the prelude, verse, pre-chorus, chorus, interlude, bridge, coda, etc. of the audio medium.

[0068] The system first extracts multi-dimensional features such as spectral features, beat features, and harmonic features of the input specified music for music interval identification, and then uses a music structure analysis algorithm, such as a hidden Markov model or a recurrent neural network, to segment the audio. The system can identify different functional intervals of the music: Prelude: Usually located at the beginning of the music, it plays an introductory role, characterized by a simple melody and a rhythm foreshadowing; Verse: It contains the main narrative content of the song, usually with complete lyrics and a relatively stable melody. Pre-chorus: The transitional part connecting the verse and the chorus, usually with changes in mood or rhythm to set the stage for the chorus. Chorus: The climax part of the song, usually with a louder and more memorable melody and a more intense emotional expression. Interlude: A pure music section between two singing parts, usually with instrument solos or special sound effects. Bridge: Usually appears in the second half of the song, with a distinct melody and harmony progression different from the verse and the chorus. Outro: The ending part of the song, usually a repetition of the theme or a gradual fade-out.

[0069] The system saves this interval information as time markers. For example, the intro may be from 0 to 20 seconds, the first verse is from 20 to 50 seconds, the first chorus is from 50 to 80 seconds, etc.

[0070] After completing the music interval recognition, the system performs rhythm point recognition. The system uses a beat tracking algorithm to detect the beat positions of the music and detects the strong beat positions and music turning points through energy analysis and spectral changes. The system pays special attention to positions such as the start of the chorus, the addition of drum beats, and the music climax, which are usually ideal video timing points.

[0071] The system saves the recognized music intervals and timing point position information in a structured manner, providing an accurate music timeline reference for subsequent video editing.

[0072] Step S13: Segment the video medium into shots and extract all the shots of the video. In this step, the system segments the video medium separated from the video to be edited and extracts all the shots of the video medium. Among them, the section between every two scene switches is regarded as a shot.

[0073] The system adopts a scene change detection method that combines multiple features, such as color histogram difference, edge feature change, and motion vector analysis, to improve the accuracy of scene change detection. The system also uses a machine learning model to verify potential scene change points, reducing false detections and missed detections.

[0074] For gradual scene changes (such as fade-in, fade-out, dissolve, etc.), the system adopts a special detection algorithm to identify such smooth transition scene changes by analyzing the change trend of consecutive multiple frames.

[0075] The system takes the detected scene change points as the boundaries of the shots, and the video segments between every two adjacent change points are marked as independent shots. The system assigns a unique ID to each shot and records basic information such as its start time, end time, and duration.

[0076] In this way, the original video is segmented into multiple independent shot segments, forming a video material library, which provides a basis for subsequent shot selection and assembly.

[0077] Step S14: Perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the shot motion data information and subtitle information. In this step, the system performs frame extraction analysis on the video medium, extracts each frame of the video, and calculates the shot motion data information and subtitle information.

[0078] The system adopts an adaptive frame extraction strategy, increasing the frame extraction frequency for segments with intense motion and decreasing the frame extraction frequency for static scenes to balance the calculation efficiency and analysis accuracy.

[0079] For the calculation of the shot motion data information, the system uses high-precision optical flow algorithms, such as the Farneback optical flow or deep learning optical flow models such as FlowNet, to calculate the pixel-level motion vectors between adjacent frames. The system analyzes the statistical characteristics of the optical flow field and extracts various motion features: Average motion amplitude: Reflects the overall motion intensity. Motion direction consistency: Reflects the coordination of motion. Motion complexity: Reflects the complexity of the motion pattern. Local motion saliency: Detects the regions with significant motion in the frame.

[0080] The system also distinguishes between camera motion (such as panning, rotation, zooming in and out) and object motion, and assigns weighted scores to different types of motion. After comprehensive evaluation of these motion features, a motion activity score for the shot is formed, which is used for subsequent selection of high-energy shots.

[0081] For the extraction of subtitle information, the system uses advanced text detection and recognition technologies. The system first uses a text detection algorithm (such as EAST or TextBoxes++) to locate the text regions in the video frame, and then uses an OCR engine (such as Tesseract or a commercial OCR API) to recognize the text content.

[0082] The system not only extracts the text content of the subtitles, but also analyzes features such as the position, size, duration, and appearance frequency of the subtitles to evaluate the importance of the subtitles. The system also uses natural language processing technologies to perform a preliminary analysis of the subtitle content, identifying information such as keywords and sentiment tendencies, providing richer semantic information for subsequent selection of ending words.

[0083] Step S15: Separate the background sound from the audio medium and calculate the energy value of the background sound. In this step, the system processes the audio medium using more advanced sound source separation technology. The system adopts a deep learning sound source separation model, such as Demucs or Open - Unmix, to separate the audio into multiple parts, including vocals, drums, bass, and other instruments.

[0084] The system focuses on the background music part and combines the separated non - vocal parts into background sound. The system conducts multi - dimensional energy analysis on the background sound: Short - term energy: The background sound is segmented into small time windows (usually 20 - 50 milliseconds), and the energy value of each window is calculated; Band energy: The audio signal is decomposed into different frequency bands to analyze the energy distribution of low - frequency, mid - frequency, and high - frequency; Rhythm energy: Analyze the energy change pattern synchronized with the beat; Spectral flux: Calculate the inter - frame change rate of the spectrum, reflecting the dynamic changes of the music.

[0085] The system combines these energy features to construct the energy curve of the background sound and identify the high - energy peak regions. The system further aligns the energy curve with the previously identified music structure to determine the energy characteristics of each music section (such as chorus, bridge, etc.).

[0086] The system marks the high - energy regions as "high - energy intervals", which usually correspond to the climax parts or parts with strong rhythms of the music and are suitable for matching high - impact and exciting shots.

[0087] Step S16: Perform mixed - cut video assembly: Based on the music section recognition results, calculate the end time of the mixed - cut video; Select a specified shot from the extracted shots as the intro; Sort based on the calculated shot movement data information, and select exciting shots located in the high - energy intervals of the sound as the middle segments; Select the ending words based on the subtitle information and find the corresponding shots as the ending; Perform video assembly and add the video logo to generate the video clip result.

[0088] In this embodiment, the system performs more refined mixed - cut video assembly work based on various analysis results obtained in the previous steps.

[0089] First, the system calculates the end time of the mixed - cut video based on the music section recognition results. The system analyzes the integrity of the music structure and preferentially selects the natural boundaries of the music structure (such as the end of the chorus, the start of the coda, etc.) as the end point of the video. The system also considers the target video length that the user may set and selects the music structure boundary closest to the target length as the end time while maintaining the integrity of the music structure.

[0090] Secondly, the system selects a specified shot from the extracted shots as the opening title. Based on all the shots of the extracted video, the system picks the shots with the main label of scenery as the opening title. The system uses a deep learning scene classification model, such as ResNet or EfficientNet, to classify the key frames of each shot and identify the shots containing open scenes such as natural scenery and urban landscapes.

[0091] The system stipulates the total duration of the opening title, usually 10%-15% of the total video length. The duration of each opening title shot is limited to no more than a specified time (usually 2-3 seconds), and the system will use the music beat position as the shot switching point to achieve the effect of audio-visual synchronization. The system will also exclude the shots containing subtitles to maintain the purity and visual impact of the opening title.

[0092] Next, the system sorts based on the calculated shot movement data information and selects the high-energy shots located in the high-energy interval of the sound as the middle part of the video. The system sorts all the shots from high to low according to the score of the shot movement data information. The system preferentially selects the shots with higher movement data scores, which usually contain more dynamic elements and visual impact.

[0093] The system matches the high-score shots with the previously identified high-energy intervals of the music to ensure that visually strong shots are paired with the parts where the music rhythm is strong and the energy is high. The system will also exclude the shots containing subtitles to avoid subtitle interference with the visual experience. These high-energy shots are selected as the middle part of the video, constituting the main body of the video.

[0094] Then, the system selects the ending words based on the subtitle information and finds the corresponding shots as the ending of the video. The system obtains the previously extracted subtitle information and inputs these subtitle texts into a large language model for processing. Based on semantic understanding and sentiment analysis, the large language model selects the text fragments suitable as the video ending words from the subtitles, that is, the "ending words".

[0095] The ending words usually have the characteristics of summarization, theme-pointing, or emotional sublimation. The system finds the video segments corresponding to the ending words, retrieves and intercepts the shots containing the subtitles as the ending shots of the video. If there are multiple shots containing the same or similar ending words, the system will select the shot with higher picture quality and stronger emotional expression.

[0096] Finally, the system assembles the selected shots as the opening title, the high-energy shots of the middle part of the video, and the corresponding shots as the ending part of the video according to the calculated ending time of the mixed video. The system stitches these shots together in chronological order, ensuring that the shot switching points are aligned with the music beats to achieve the effect of audio-visual synchronization.

[0097] The system will also add audio-visual effects, such as adding smooth transition effects, adjusting the color style to maintain consistency, and adding appropriate visual effects to enhance the visual experience. The system will also add a video logo at an appropriate position in the video (usually at the beginning or end) to identify the creator or brand information of the video.

[0098] Through the above steps, the system generates the final video clip result, a mixed video based on the original video material, organized according to the music rhythm, and with high visual impact.

[0099] Embodiment III An automated video editing and processing method based on sound and optical flow provided by Embodiment III of the present invention includes the following steps: Step S21: Obtain the video to be edited, separate the audio and video of the video to be edited, and obtain the audio medium and video medium respectively; In this embodiment, the system receives the original video file to be edited through the user interface or API interface. The system supports the input of multiple video formats, including but not limited to common formats such as MP4, AVI, MOV, and WMV. The system uses a professional audio-visual processing library to decode and separate the video.

[0100] The system decodes the video file into the original data stream, and then extracts the audio stream and video stream respectively. The audio stream is saved in a high-quality lossless audio format (such as WAV or FLAC) to ensure the accuracy of subsequent audio analysis; the video stream is saved as a silent video file, retaining the resolution and frame rate of the original video. The system also records the metadata information of the original video, such as duration, encoding format, creation time, etc., for subsequent processing reference.

[0101] Step S22: Identify the music intervals of the input specified music, and then identify the rhythm points to find the positions where the music makes beats; In this embodiment, the system first preprocesses the input specified music for music interval identification, including noise suppression, volume normalization, etc., to improve the accuracy of subsequent analysis. Then the system performs music interval identification and analyzes the complete music structure of the audio medium, such as the prelude, verse, pre-chorus, chorus, interlude, bridge, coda, etc.

[0102] The system uses a multi-modal feature fusion method for music structure analysis. The system extracts multi-dimensional features of the input specified music, such as time-domain features (such as energy envelope, zero-crossing rate), frequency-domain features (such as spectral centroid, Mel-frequency cepstral coefficients), and harmonic features (such as chord progression, key change). The system uses a self-supervised learning model to encode these features, and then uses a segmentation algorithm (such as adaptive clustering or recurrent neural network) to segment the audio.

[0103] The system can accurately identify different functional intervals of music and mark the start time, end time, and interval type for each interval. For example, the system may identify: the intro (0 - 18 seconds), the first verse (18 - 45 seconds), the pre-chorus (45 - 53 seconds), the first chorus (53 - 80 seconds), etc.

[0104] After completing the music interval recognition, the system performs rhythm point recognition. The system first uses a beat tracking algorithm to detect the basic beat positions of the music, and then finds the time points suitable for video beat matching through multi-feature analysis. The system focuses on the following types of potential beat matching points: The positions of strong beats, especially the first beat of each measure; The boundary points of the music structure, such as the position where the chorus starts; The points of sudden volume change, such as the position where the drumbeat is added or the music bursts; The positions with significant spectral changes, such as the position where a new instrument is added or the timbre changes.

[0105] The system assigns an importance score to each potential beat matching point and saves them as a beat matching point list in chronological order to provide an accurate music timeline reference for subsequent video editing.

[0106] Step S23: Segment the video medium into shots and extract all the shots of the video; In this step, the system segments the video medium separated from the video to be edited and extracts all the shots of the video medium. Among them, the section between every two scene switches is regarded as a shot.

[0107] The system uses a deep learning scene change detection method and analyzes the visual differences between adjacent frames using a pre-trained convolutional neural network model. The system can not only detect scene changes of the hard cut (direct switch) type, but also accurately identify various smooth transition scene changes, such as fade-in, fade-out, dissolve, wipe, push-pull and other special effect switches.

[0108] The system post-processes the detected scene change points, including removing mis-detected points (such as mis-detections caused by flashes and fast movements) and supplementing missed detected points (such as scene changes in low-contrast scenes). The system also analyzes the type and characteristics of the scene changes to provide a reference for subsequent video assembly.

[0109] The system marks the video segment between every two adjacent change points as an independent shot and assigns a unique ID to each shot. The system also extracts the key frame of each shot, usually choosing the middle frame of the shot or the frame with the richest visual information as the key frame for subsequent content analysis.

[0110] The system creates a detailed metadata file for each shot, including start time, end time, duration, scene transition type, key frame information, etc. These shots form a video material library, providing a basis for subsequent shot selection and assembly.

[0111] Step S24: Conduct frame extraction analysis on the video medium, extract each frame of the video, and calculate shot motion data information and subtitle information; In this step, the system conducts intelligent frame extraction analysis on the video medium, extracts each frame of the video, and calculates shot motion data information and subtitle information.

[0112] The system adopts a content-aware frame extraction strategy, increasing the frame extraction frequency for segments with fast-changing visual content and decreasing the frame extraction frequency for static scenes. The system usually conducts frame extraction at a frequency of 5 - 10 frames per second, but will dynamically adjust according to the content complexity.

[0113] For the calculation of shot motion data information, the system uses a high-precision optical flow algorithm combined with a deep learning model for analysis. The system first calculates the dense optical flow field between adjacent extracted frames, and then extracts various motion features: Global motion features: Reflect camera motion, such as translation, rotation, zooming in and out, swaying, etc.; Local motion features: Reflect the motion of objects in the frame, such as human actions, object movements, etc.; Motion consistency: Evaluate the degree of consistency of motion direction and amplitude; Motion complexity: Evaluate the complexity of motion patterns; Motion rhythm: Evaluate the temporal variation pattern of motion.

[0114] The system comprehensively evaluates these features and calculates a motion activity score for each shot. Shots with high scores usually contain more dynamic elements and visual impact, and are suitable as high-energy segments.

[0115] For the extraction of subtitle information, the system uses a multi-stage text detection and recognition process. The system first uses a deep learning text detection model to locate the text regions in the video frame, and then uses a text recognition model to convert the text in the image into editable text information.

[0116] The system not only extracts the text content of the subtitles, but also analyzes features such as the position, size, font, color, appearance time, and duration of the subtitles. The system also uses natural language processing technology to conduct semantic analysis on the subtitle content, identifying information such as keywords, entities, sentiment tendencies, and semantic importance, providing rich semantic support for subsequent end word selection.

[0117] Step S25: Separate the background sound from the audio medium and calculate the energy value of the background sound; In this step, the system uses advanced sound source separation technology to finely process the audio medium. The system adopts a deep learning sound source separation model, which can separate the mixed audio into multiple independent parts such as vocals, drumbeats, bass, and other instruments.

[0118] The system focuses on the background music part and combines the separated non-vocal parts into background sound. The system conducts multi-dimensional energy analysis and feature extraction on the background sound: Time-domain energy analysis: Calculate the short-time energy curve to reflect the volume change; Frequency-domain energy analysis: Calculate the energy distribution in different frequency bands (low frequency, medium frequency, high frequency); Rhythm energy analysis: Analyze the energy change pattern synchronized with the beats; Spectral flux analysis: Calculate the inter-frame change rate of the spectrum to reflect the dynamic changes of the music; Timbre feature analysis: Extract timbre-related features such as spectral centroid and spectral flatness.

[0119] The system combines these features to construct a multi-dimensional energy feature curve of the background sound and identifies the high-energy peak regions. The system further aligns the energy features with the previously identified music structure to determine the energy characteristics of each music section (such as chorus, bridge, etc.).

[0120] The system marks the high-energy regions as "high-energy intervals", which usually correspond to the climax parts or parts with strong rhythms of the music. The system also calculates the energy change rate and identifies the intervals where the energy rises rapidly. These intervals are usually emotional turning points and are suitable for arranging shot transitions with strong visual impact.

[0121] Step S26: Assemble the mixed-cut video: Based on the music section recognition result, calculate the end time of the mixed-cut video; Select a specified shot from the extracted shots as the opening title; Sort based on the calculated shot movement data information, and select the high-energy and exciting shots located in the high-energy intervals of the sound as the middle segments; Select the ending words based on the subtitle information and find the corresponding shots as the ending; Assemble the video and add the video logo to generate the video editing result.

[0122] In this embodiment, the system performs refined mixed-cut video assembly work based on various analysis results obtained in the previous steps.

[0123] First, the system calculates the end time of the mixed-cut video based on the music section recognition result. The system analyzes the integrity of the music structure and the emotional arc, and preferentially selects the natural boundaries of the music structure (such as the end of the chorus, the start of the coda, etc.) as the end point of the video. The system also considers the target video length that the user may set and the platform requirements (such as the duration limit of short video platforms), and selects the most suitable music structure boundary as the end time while maintaining the integrity of the music structure.

[0124] Secondly, the system selects a designated shot from the extracted shots as the title. Based on all the shots of the extracted video, the system selects shots with the subject label as scenery as the title. The system uses deep learning scene classification and content understanding models to analyze the content of each shot and identify shots containing open scenes such as natural scenery and urban landscapes.

[0125] The system agrees on the total length of the title, which is usually 10%-15% of the total length of the video. The length of each title shot is limited to a specified time (usually 2-3 seconds), and the system will use the music card position as the shot switching point to achieve the effect of audio and video synchronization. The system will also remove shots containing subtitles to maintain the purity and visual impact of the title. The system will also consider the visual coherence between shots, and select a combination of shots with similar tones, compositions or themes to enhance the overall beauty of the title.

[0126] Next, the system sorts the shots based on the calculated shot motion data information and selects the high-energy shots in the high-energy sound range as the clips in the film. The system sorts all shots from high to low according to the scores of the shot motion data information. The system gives priority to shots with higher motion data scores, which usually contain more dynamic elements and visual impact.

[0127] The system matches high-scoring shots with previously identified high-energy sections of the music, ensuring that the parts with strong music rhythms and high energy are matched with shots with equally strong visual effects. The system also removes shots containing subtitles to prevent subtitles from interfering with the visual experience. The system also considers the diversity of shot content and rhythm changes to avoid using overly similar shots in succession to maintain visual freshness.

[0128] Then, the system selects the ending words based on the subtitle information and finds the corresponding shots as the ending. The system obtains the previously extracted subtitle information and inputs these subtitle texts into the large language model for processing. The large language model selects text fragments from the subtitles that are suitable as the video ending, namely the "ending words", based on semantic understanding and sentiment analysis.

[0129] Ending words usually have the characteristics of summary, theme or emotional sublimation. The system finds the video clip corresponding to the ending words, retrieves and captures the shots containing the subtitles as the ending shots. If there are multiple shots containing the same or similar ending words, the system will select the shots with higher picture quality, stronger emotional expression and better composition. The system will also consider the consistency of the ending words with the overall video theme, and select the ending words that can echo the beginning or theme of the video to form a complete narrative closed loop.

[0130] Finally, based on the calculated end time of the mixed video, the system assembles the selected shots as the opening, the exciting shots in the middle part of the video, and the corresponding shots as the ending part. The system stitches these shots together in chronological order, ensuring that the shot transition points are aligned with the music beats to achieve the effect of audio-visual synchronization.

[0131] In the embodiments of the present invention, audio and video effects are also added, such as adding smooth transition effects (such as fade-in, fade-out, dissolve, slide, etc.), adjusting the color style to maintain consistency (such as applying a unified color filter or tone), and adding appropriate visual effects (such as flicker, blur, zoom, etc.) to enhance the visual experience. The system also adds a video logo at an appropriate position in the video (usually at the beginning or end) to identify the creator or brand information of the video. The system may also add subtitles, watermarks, or other brand elements to enhance the professionalism and recognition of the video.

[0132] Through the above steps, the system generates the final video editing result, a mixed video based on the original video materials, organized according to the music rhythm, and with high visual impact and emotional expressiveness. And the present invention also has the following technical effects: The present invention can realize the joint modeling of the dynamic association between sound and picture, avoiding the problem of audio-visual disconnection. Especially when video editing according to the music rhythm points is required, the present invention can accurately identify the music intervals and rhythm points, and can achieve high-quality audio-visual synchronization.

[0133] Secondly, the present invention can accurately capture the exciting shots in the video and the high-energy intervals of the music, improving the editing effect.

[0134] Thirdly, the present invention achieves a balance between efficiency and effect. Through intelligent editing, it can be efficiently completed in aspects such as shot segmentation, motion data calculation, and subtitle information extraction.

[0135] Therefore, the automated video editing processing method based on sound and optical flow in the embodiments of the present invention realizes the joint analysis and processing of audio-visual information, improving the automation degree and editing effect of video editing.

[0136] It should be noted that Embodiment 1, Embodiment 2, and Embodiment 3 are all a kind of automated video editing processing method based on sound and optical flow.

[0137] The following further elaborates on the present invention through specific application embodiments: Embodiment 4 As Figure 3 shown, an automated video editing processing method based on sound and optical flow provided by this specific application embodiment 4 includes the following steps: S40. Obtain the video to be edited, separate the audio and video of the video to be edited, and obtain the audio medium and video medium respectively; S41. Analyze the specified music input (intro, verse, pre-chorus, chorus, interlude, bridge, outro), then identify the beat points, and find the positions where the music can be timed (down-beat points); The analysis of the audio medium can be completed in one minute (1min); Among them, intro: the part at the beginning of the audio, usually used to introduce the audio theme and set the atmosphere.

[0138] Verse: the main narrative part of the audio, usually containing the story or emotional core of the audio, and usually repeated multiple times.

[0139] Pre-chorus: the part connecting the verse and the chorus, used to increase the tension and guide the listener into the chorus.

[0140] Chorus: the most core and memorable part of the audio, with repeated lyrics, conveying the main theme or emotion.

[0141] Interlude: the instrumental part between the chorus or verse, providing a break and adding layers to the music.

[0142] Bridge: a contrasting part, usually appearing once in the song, providing a different melody or emotion, highlighting the changes in the song.

[0143] Outro: the ending part of the song, used to end the whole song, usually reviewing or reaffirming the theme.

[0144] S42. Segment the video medium into shots and extract all the shots of the video; Segment the video medium into shots. Segmenting the shots means analyzing the whole film and extracting all the shots of the film. The section between every two scene transitions is regarded as a shot. The process of segmenting the shots can be completed in 10 minutes (minutes); S43. Perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the lens movement data information and subtitle information; Steps S42 and S43 are to extract frames from the film and analyze them, extract each frame of the video to calculate the required data. The required data includes lens movement data information (the faster the lens switches, the faster the lens moves, and the more active the movement data), and subtitle information (providing data support for the large model to select the ending words).

[0145] S44. Separate the background sound from the audio medium and calculate the energy value of the background sound; In the embodiment of this step, the audio of the video is separated, the background sound is analyzed, and the energy of the background music is calculated (the louder and more miscellaneous the background music is, the higher the music energy).

[0146] S45. Assemble the mixed video: Based on the music interval recognition result, calculate the end time of the mixed video; select a specified shot from the extracted shots as the title; sort based on the calculated shot movement data information, and select the high-energy shots in the high-energy sound interval as the middle clips; select the ending words based on the subtitle information, find the corresponding shots as the ending; perform video assembly and add a video logo to generate the video editing result.

[0147] In the embodiment of the present invention, when assembling the mixed video, the following is specifically executed: S451. Calculate the end time of the mixed video: Based on the music interval recognition result; S452. Select the title shot: Based on the shot analysis result, select the shot with the main body label of scenery as the title. The total duration of this title is 10s, and the time limit for each shot does not exceed 2s, and basically realizes music timing (that is, the music timing is the shot switching point), and the shots with subtitles are excluded; S453. Select the high-energy middle shots: Select based on the high-energy sound interval, sort according to the motion calculation data, screen out the most intense part of the motion within the shot, and exclude the shots with subtitles; S454. Select the ending shot: Based on the ending words selected by the large language model, find the corresponding segment of the ending words, and retrieve and intercept the corresponding shot; S455. Video assembly, add audio and video effects: Add volume balance, fade-in and fade-out effects for audio and video; S456. Add the video logo: Add the logo at the end of the video.

[0148] In this specific embodiment, obtain the video to be edited: that is, select the original video file that the user wants to edit.

[0149] Audio-video separation: Separate the audio and video content in the video for separate processing. Two media are obtained: an audio medium (containing sound) and a video medium (containing images).

[0150] Then perform music interval and rhythm point recognition: Among them, music interval recognition: Analyze the specified music (that is, the background music selected for video editing) to identify different music segments (such as introduction, main melody, climax, etc.).

[0151] And rhythm point recognition: Find the key rhythm points (such as beats) in the music, which will help arrange the timing in the video.

[0152] Regarding video segmentation: The present invention decomposes the video medium into multiple shots and extracts the content of each shot, which makes subsequent editing more flexible.

[0153] Frame extraction and analysis: It extracts each frame from the video and analyzes the motion information and subtitle information of each frame to obtain more editing materials and context information.

[0154] Regarding the extraction and energy calculation of the original background sound of the video to be edited: Separate the background sound: Extract the background sound (such as ambient sound) from the originally separated audio medium and calculate its energy value to analyze which areas can be used for editing or enhancing the effect.

[0155] Then perform mixed video assembly: Calculate the end time: Based on the result of music interval recognition, determine the end time of the mixed video. Regarding the selection of the opening shot: Select a suitable shot from the extracted shots as the beginning of the video.

[0156] Regarding sorting and selecting high-energy shots: Specifically, according to the shot motion data, select the shots that perform outstandingly in the high-energy areas of the music (such as the climax part) as the main content of the video.

[0157] Ending selection and video assembly: Find the appropriate ending words according to the subtitle information, select the corresponding shots, and finally complete the video assembly and add the LOGO to generate the editing result.

[0158] Thus, the present invention has the following effects: 1) High-efficiency automation: Through automated analysis, recognition, and editing, it saves a large amount of time and effort required for manual editing and improves work efficiency.

[0159] 2) Precise editing: Using the rhythm and high-energy intervals of the music for editing can make the finally generated video more professional and attractive, and perfectly combine with the music and rhythm.

[0160] 3) Rich content: Through frame extraction and analysis and shot extraction, it ensures that the video contains sufficient diversity and information, making the viewing experience of the audience more vivid.

[0161] 4) Background sound control: The analysis of the background sound and the calculation of the energy value help to ensure the balance of the audio during the editing process and enhance the overall music effect.

[0162] 5) Professional output: The finally generated video is not only content-attractive but also has professional elements such as a LOGO, improving brand recognition.

[0163] As can be seen from the above, the embodiments of the present invention provide an automated video editing method that can deeply integrate audio-visual features, is lightweight and has strong dynamic adaptability. The automated video editing based on sound and optical flow of the present invention generates high-quality short videos, attracts users' attention, and improves the efficiency and quality of video generation.

[0164] Exemplary device As Figure 4 shown, the embodiments of the present invention provide an automated video editing and processing device based on sound and optical flow. The device includes: An audio-video separation module 310, configured to obtain a video to be edited, perform audio-video separation on the video to be edited, and respectively obtain an audio medium and a video medium; A music interval recognition module 320, configured to perform music interval recognition on the input specified music, and then perform rhythm point recognition to find the positions where the music hits the beats; A lens segmentation module 330, configured to perform lens segmentation on the video medium and extract all the lenses of the video; A frame extraction and analysis module 340, configured to perform frame extraction and analysis on the video medium, extract each frame of the video, and calculate the lens motion data information and subtitle information; A background sound energy value calculation module 350, configured to separate the background sound from the audio medium and calculate the energy value of the background sound; A mixed video assembly module 360, configured to perform mixed video assembly: based on the music interval recognition result, calculate the end time of the mixed video; select a specified lens from the extracted lenses as the start of the video; perform sorting based on the calculated lens motion data information, and select high-energy and exciting lenses located in the high-energy interval of the sound as the middle segments; select the ending words based on the subtitle information, find the corresponding lenses as the end of the video; perform video assembly and add a video logo to generate a video editing result, as described above.

[0165] Based on the above embodiments, the present invention also provides an intelligent terminal, and its principle block diagram can be as Figure 5 shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an automated video editing and processing method based on sound and optical flow. The database of the intelligent terminal is used to store the automated video editing and processing program based on sound and optical flow.

[0166] Those skilled in the art can understand that Figure 5 The block diagram of the principle shown in Figure 5 is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0167] In one embodiment, an intelligent terminal is provided, which includes a memory and one or more programs. One or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain the video to be clipped, perform audio-video separation on the video to be clipped to obtain an audio medium and a video medium respectively; perform music interval recognition on the input specified music, then perform rhythm point recognition to find the positions where the music makes beats; perform shot segmentation on the video medium to extract all the shots of the video; perform frame extraction analysis on the video medium, extract each frame of the video and calculate the shot movement data information and subtitle information; separate the background sound from the audio medium and calculate the energy value of the background sound; perform mixed video assembly: based on the music interval recognition result, calculate the end time of the mixed video; select a specified shot from the extracted shots as the title; sort based on the calculated shot movement data information and select high-energy shots located in the high-energy interval of the sound as the middle segments; select the ending words based on the subtitle information and find the corresponding shots as the end; perform video assembly and add a video logo to generate the video clipping result.

[0168] Preferably, the music interval analysis includes analyzing the prelude, verse, pre-chorus, chorus, interlude, bridge, and coda of the audio medium.

[0169] Further, the step of performing shot segmentation on the video medium to extract all the shots of the video includes: performing shot segmentation on the video medium separated from the video to be clipped, and extracting all the shots of the video medium, where the section between every two scene switches is regarded as a shot.

[0170] Further, the step of performing frame extraction analysis on the video medium, extracting each frame of the video and calculating the shot movement data information and subtitle information includes: performing frame extraction analysis on the video medium, extracting each frame of the video and calculating the shot movement data information and subtitle information, where the shot movement data information is the active movement data shown during the shot switching, and the subtitle information provides data support for the large model to select the ending words.

[0171] Further, the step of selecting a specified shot from the extracted shots as the opening title includes: based on all the shots of the extracted video, selecting the shots with the subject label of scenery as the opening title; agreeing on the total duration of the opening title, limiting the time of each shot not to exceed the specified time, and implementing the music beat as the shot transition point, and excluding the shots with subtitles.

[0172] Further, the step of sorting based on the calculated shot motion data information and selecting the exciting shots located in the high-energy interval of the sound as the middle part of the video includes: based on the calculated shot motion data information, sorting the shot motion data information in descending order; according to the sorting of the shot motion data information, screening out the exciting shots with the sorted motion data information in the front, and excluding the shots with subtitles as the exciting shots in the middle part of the video.

[0173] Further, the step of selecting the ending word based on the subtitle information and finding the corresponding shot as the ending of the video includes: obtaining the extracted subtitle information; inputting the subtitle information into a large language model for processing, and selecting the ending word based on the large language model; and finding the corresponding segment to the ending word, retrieving and intercepting the corresponding shot as the ending shot of the video; the step of assembling the video and adding the video logo to generate the video editing result includes: according to the calculated ending time of the mixed video, assembling the selected shots as the opening title, the exciting shots in the middle part of the video, and the corresponding shots as the ending part of the video, and adding the audio-video effects and adding the video logo to generate the video editing result, as described above.

[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0175] In summary, the present invention provides an automated video editing processing method, device, intelligent terminal, and storage medium based on sound and optical flow. By identifying the music intervals and rhythm points of the input specified music and combining the lens segmentation and motion data information analysis of the video medium, the present invention realizes the joint modeling of the dynamic association between sound and picture, and solves the problem of audio-visual separation in traditional methods; by performing frame extraction analysis on the video medium to calculate the lens motion data information and calculating the background sound energy value of the audio medium, the dynamic adaptability is improved, and the optical flow calculation in complex motion scenes and the rhythm extraction of non-periodic sounds can be effectively processed; the method of the present invention does not rely on high-computing-power GPU real-time inference, and balances efficiency and effect through preprocessing and intelligent assembly, and can be adapted to mobile terminals or low-configuration devices; through precise audio-visual synchronization and dynamic editing, high-quality short videos can be generated, effectively attracting users' attention.

Claims

1. An automated video editing processing method based on sound and optical flow, characterized in that, include: Obtaining a video to be edited, and performing audio and video separation on the video to be edited to obtain an audio medium and a video medium respectively; Identify the music interval of the input specified music, then identify the rhythm point, and find the location of the music card point; Segmenting the video medium into shots to extract all shots of the video; Performing frame extraction analysis on the video medium, extracting each frame of the video to calculate lens motion data information and subtitle information; Separating background sound from the audio medium and calculating energy value of the background sound; Assemble the mixed-cut video: calculate the end time of the mixed-cut video based on the music interval recognition result; select the specified shot from the extracted shots as the title; sort the calculated shot motion data information and select the high-energy shot in the high-energy sound interval as the mid-film segment; select the end word based on the subtitle information and find the corresponding shot as the end; assemble the video and add the video logo to generate the video editing result.

2. The automated video editing processing method based on sound and optical flow according to claim 1, wherein The music interval analysis includes analyzing the prelude, verse, introduction, chorus, interlude, bridge and outro of the audio medium.

3. The automated video clip processing method based on sound and optical flow according to claim 1, wherein, The step of segmenting the video medium into shots and extracting all the shots of the video comprises: Shot segmentation is performed on the video medium separated from the video to be edited, and all shots of the video medium are extracted, wherein the segment between every two scene switches is regarded as one shot.

4. The automated video clip processing method based on sound and optical flow according to claim 1, characterized in that, The step of performing frame extraction analysis on the video medium and extracting each frame of the video to calculate the lens motion data information and subtitle information comprises: The video medium is subjected to frame extraction analysis, and each frame of the video is extracted to calculate the lens motion data information and subtitle information, wherein the lens motion data information is the active motion data exhibited by the lens during the switching process, and the subtitle information is used to provide data support for selecting the end word for the large model.

5. The automated video clip processing method based on sound and optical flow according to claim 1, wherein The step of selecting a designated shot from the extracted shots as the title sequence comprises: Based on all the shots of the extracted video, the shots with the subject label of scenery are selected as the opening. The total length of the opening is agreed upon, the time limit for each shot does not exceed the specified time, and the music card point is used as the shot switching point, and the shots with subtitles are eliminated.

6. The automated video clip processing method based on sound and optical flow according to claim 5, wherein The step of sorting the calculated shot motion data information and selecting the high-energy shots in the high-energy sound range as the clips in the film includes: Based on the calculated lens motion data information, the lens motion data information is sorted in descending order; According to the sorting of the lens motion data information, the high-energy shots with the motion data information sorted in the front are selected, and the shots with subtitles are eliminated, as the high-energy shots in the film clips.

7. The automated video clip processing method based on sound and optical flow according to claim 6, characterized in that The step of selecting the ending word based on the subtitle information and finding the corresponding shot as the ending includes: Get the extracted subtitle information; Input the subtitle information into the large language model for processing, select an ending word based on the large language model; find a segment corresponding to the ending word, retrieve and capture the corresponding shot as the ending shot; The steps of assembling the video and adding the video logo to generate the video editing result include: According to the calculated end time of the mixed video, assemble the selected shots as the opening, the exciting shots of the video clips in the middle, and the corresponding shots of the ending segment, and add audio-visual effects and the video logo to generate the video editing result.

8. An automated video editing processing device based on sound and optical flow, characterized in that, The device includes: An audio-visual separation module, configured to obtain the video to be edited, separate the audio and video of the video to be edited, and obtain the audio medium and the video medium respectively; A music interval recognition module, configured to recognize the music interval of the input specified music, then recognize the rhythm points, and find the positions where the music makes a beat; A shot segmentation module, configured to segment the video medium and extract all the shots of the video; A frame extraction analysis module, configured to perform frame extraction analysis on the video medium, extract each frame of the video, and calculate the shot movement data information and the subtitle information; A background sound energy value calculation module, configured to separate the background sound from the audio medium and calculate the energy value of the background sound; A mixed video assembly module, configured to assemble the mixed video: based on the music interval recognition result, calculate the end time of the mixed video; select the specified shot from the extracted shots as the opening; based on the calculated shot movement data information, sort them, and select the exciting shots in the high-energy sound interval as the middle video clips; select the ending words based on the subtitle information, find the corresponding shots as the ending; perform video assembly and add the video logo to generate the video editing result.

9. An intelligent terminal, characterized in that, It includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the method according to any one of claims 1-7.