Face-based automatic face score video editing processing method and device, and terminal
Through lens sharding and facial attribute analysis combined with music beat synchronization, the problem of insufficient face screening and audio and video synchronization in the existing technology is solved, and high-value and smooth video editing is achieved, which is suitable for short video creation.
Patent Information
- Application Number
- CN202510586851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video editing technology cannot effectively combine face screening with audio and video synchronization, resulting in missing appearance lenses or misalignment of rhythm, insufficient dynamic adaptability, imbalance in editing efficiency and effect, especially in complex sports scenarios, it is difficult to achieve efficient automated editing.
Through lens sharding, face attribute analysis and music beat synchronization, an audio-visual collaborative editing framework is built to achieve accurate screening of target faces and music rhythm matching, and combine multi-dimensional face attribute calculation and music rhythm recognition to generate high-value and smooth short videos.
It realizes efficient automation of target face video editing, ensures that the picture and audio are synchronized, improves the quality and fluency of video editing, meets the needs of creators, and is suitable for lightweight processing on mobile.
Smart Images

Figure CN120302106A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly relates to a method, device, intelligent terminal and storage medium for automated video editing and processing based on facial beauty Background Art
[0002] With the rapid development of the short video industry, the market size exceeded one trillion yuan in 2024, and users spent more than 2 hours watching videos per day on average. Creators' demand for efficient and automated editing tools has increased sharply. Current video automated editing technologies mainly include key frame detection, scene segmentation, and audio beat-based synchronization technology. Key frame detection identifies significantly changing frames by analyzing the differences between frames for editing reference, such as at the scene transition of a landscape video; scene segmentation uses image algorithms to divide video scene segments, like distinguishing different shooting locations in a movie; audio beat synchronization matches the video transitions with the audio beats and is commonly used in music short videos. However, in the existing technologies for video editing, there is basically no good automated video editing based on a specified person's facial beauty. At the same time, there are serious problems of audio-visual disconnection in the existing video editing technologies, and traditional methods process audio and video signals separately. For example, cutting the video according to the beat is likely to cause discontinuous actions, and the optical flow method for scene segmentation ignores the impact of sound emotion on the editing rhythm, resulting in the lack of a sense of integrity of the video. There is insufficient dynamic adaptability. In complex motion scenes, the calculation deviation of the optical flow algorithm is large. For example, in a sports event with fast and complex actions of athletes, it is easy to misalign the selection of editing points; the audio analysis module has weak ability to extract the rhythm of non-periodic sounds (ambient sounds, conversations). There is an imbalance between efficiency and effect. Although the multi-modal scheme based on deep learning can improve the effect, it relies on high-computing-power GPUs for real-time inference and is difficult to run on mobile devices or low-configuration devices, which limits the implementation of the technology.
[0003] Therefore, there is an urgent need for a method for automated video editing and processing based on facial beauty that can deeply integrate audio-visual features, is lightweight and has strong dynamic adaptability to generate high-quality short videos to meet the needs of creators and the market. In view of the above problems, the existing technologies need to be improved. Summary of the Invention
[0004] In view of the above technical problems existing in the prior art, the present invention provides a method, device, intelligent terminal and storage medium for automated video editing and processing based on facial beauty. The present invention provides a method for automated video editing and processing based on facial beauty that can deeply integrate audio-visual features, is lightweight and has strong dynamic adaptability, can automatically edit and generate high-quality short videos with the target face, meet the needs of creators and the market, and improve the efficiency and quality of video editing.
[0005] The technical solutions adopted by the present invention to solve the problems are as follows: An automated face-based video editing and processing method for appearance includes: obtaining the video to be edited and the editing requirements including the target face, segmenting the video to be edited into multiple shots; according to the editing requirements including the target face, performing frame extraction analysis on each segmented shot to analyze whether each frame includes the attribute data of the target face; and extracting the frames including the attribute data of the target face; screening the extracted frames including the attribute data of the target face, screening out the frames whose attribute data of the target face meet the editing requirements, and converting them into the shots of the filtered target face; obtaining the input specified music as the background music, identifying the music interval of the specified music, and then performing rhythm point identification to find the positions where the music makes beats; removing the shots with lines and repeated shots from the screened shots of the target face to obtain the filtered shots of the target face; splicing all the filtered shots of the target face, splicing the filtered shots of the target face according to the music beat positions, and editing the spliced video; taking the edited video as the center, expanding a predetermined duration outward on both sides, and selecting a segment with lines that match the current scene and include the target face from the cropped video through a specified large model as the final ending to generate the final edited video for output.
[0006] Further, the present application also proposes that the step of segmenting the video to be edited into multiple shots includes: performing an overall analysis on the video to be edited, extracting all the shots of the video to be edited, where the section between every two scene switches is regarded as a shot.
[0007] Further, the present application also proposes that the step of performing frame extraction analysis on each segmented shot according to the editing requirements including the target face, analyzing whether each frame includes the attribute data of the target face, and extracting the frames including the attribute data of the target face includes: where the attribute data of the target face are: the target face, and the size data, beauty degree data, and optical flow value data of the target face; according to the editing requirements including the target face, performing frame extraction analysis on each segmented shot, analyzing whether each frame includes the target face, and extracting the frames including the target face; analyzing and calculating the size data, beauty degree data, and optical flow value data of the target face for the extracted frames including the target face to obtain the frames including the attribute data of the target face.
[0008] Further, the present application also proposes a step of screening the frames containing the attribute data of the target face, screening out the frames whose attribute data of the target face meet the editing requirements, and converting them into the shots of the screened target face, which includes: screening the frames containing the attribute data of the target face, and comparing the size data, beauty degree data, and optical flow value data of the target face in the frames containing the target face extracted; with the standard beauty value size data, standard beauty value beauty degree data, and standard beauty value optical flow value data of the target face required by the editing requirements respectively; filtering out the frames that do not meet the standard beauty value size data, standard beauty value beauty degree data, and standard beauty value optical flow value data of the target face required by the editing requirements, and screening out the frames whose attribute data of the target face meet the editing requirements; and converting the frames whose attribute data of the target face meet the editing requirements into the shots of the screened target face.
[0009] Further, the present application also proposes a step of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the position where the music makes a beat, which includes: obtaining the input specified music as the background music, and performing music interval recognition on the specified music to identify the prelude, verse, pre-chorus, chorus, interlude, bridge, and coda of the specified music; then performing rhythm point recognition to find the down-beat point where the music makes a beat.
[0010] Further, the present application also proposes a step of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the position where the music makes a beat, which also includes: Method for selecting the high-energy curve of the background music: Analyze the trailer background music and the video to be edited, respectively obtain the high-energy curves of the trailer background music and the original video to be edited, and delimit various rhythm intervals; The method for selecting the high-energy curve of the trailer background music includes: performing time-domain or frequency-domain analysis on the background music to obtain the time corresponding to the chorus part of the background music, so as to obtain the high-energy curve; obtaining the high-energy curve of the background music according to the sound energy of the background music; and obtaining the high-energy curve based on the model method.
[0011] Furthermore, the present application also proposes the steps of splicing the shots of all filtered target faces, splicing the shots of the filtered target faces according to the music beat positions, and editing the spliced video, which include: obtaining the filtered target face shots, organizing the filtered target face shots, and classifying them according to chronological order or emotional expression; obtaining the identified music beat positions, marking the down-beats and the music beat positions of the climax parts according to the music rhythm; splicing the organized target face shots on the appropriate beats according to the previously identified music beat positions to ensure that each shot transition is synchronized with the music rhythm; editing the spliced video, automatically adjusting the time, special effects, and transition effects of the shot transitions to ensure the smoothness and harmony of the entire video.
[0012] Furthermore, the present application also proposes an automated video editing processing device based on facial appearance, which includes: a shot segmentation module for obtaining the video to be edited and the editing requirements including the target face, and segmenting the video to be edited into multiple shots; a target face frame extraction module for performing frame extraction analysis on each of the segmented shots according to the editing requirements including the target face, analyzing whether each frame includes the attribute data of the target face, and extracting the frames including the attribute data of the target face; a target face frame screening module for screening the extracted frames including the attribute data of the target face, screening out the frames whose attribute data of the target face meet the editing requirements, and converting them into the filtered target face shots; a music beat recognition module for obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the positions where the music beats; a target face shot filtering module for removing the shots with lines and repeated shots from the filtered target face shots to obtain the filtered target face shots; a target face shot splicing module for splicing all the filtered target face shots, splicing the filtered target face shots according to the music beat positions, and editing the spliced video; a video extension and output module for expanding a predetermined duration outward from both sides centered on the edited video, and selecting a segment of the line that conforms to the current scene and includes the target face from the cropped video through a specified large model as the final ending to generate and output the final edited video.
[0013] Furthermore, the present application also proposes an intelligent terminal, which includes a memory and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include those for executing the above methods.
[0014] Furthermore, the present application also proposes a computer-readable storage medium. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the above method.
[0015] As can be seen from the above, a face-based automated beauty video editing and processing method, apparatus, intelligent terminal, and storage medium provided by the present application achieve video editing with a high degree of matching between face beauty and music rhythm through techniques such as lens segmentation, target face attribute analysis, music beat synchronization, and intelligent expansion. It solves the problems of serious audio-visual fragmentation and insufficient dynamic adaptability in traditional methods, and has the effects of improving the editing fluency and enhancing the video expressiveness; it can automatically edit and generate high-quality short videos with the target face, meeting the needs of creators and the market, and improving the video editing efficiency and quality; the present invention can also improve the accuracy of video editing. The generated video images correspond to the high-tide parts of the audio, achieving the purpose of attracting users to watch more. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of a face-based automated beauty video editing and processing method provided in Embodiment 1 of the present invention.
[0018] Figure 2 It is a schematic diagram of the high-energy video interval structure of a face-based automated beauty video editing and processing method provided in Embodiment 2 of the present invention.
[0019] Figure 3 It is a principle block diagram of an embodiment of a face-based automated beauty video editing and processing apparatus provided by the present invention.
[0020] Figure 4 It is a schematic diagram of the internal structure principle of an intelligent terminal provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the purpose, technical solutions, and advantages of the present invention clearer and more definite, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0023] In the prior art, the rapid development of the short video industry has led to a surge in the demand of creators for efficient and automated editing tools. Traditional video editing technologies have prominent problems such as serious audio-visual fragmentation, insufficient dynamic adaptability, and imbalance between efficiency and effect. For example, in the production scenario of short videos with close-ups of people, the creator needs to quickly select fragments with people's good-looking appearances from multi-angle shot materials, and at the same time, needs to make the picture switching accurately match the rhythm of the background music. However, the prior art cannot effectively balance face screening and audio-visual synchronization, resulting in the phenomenon of missing good-looking shots or rhythm misalignment in the finished product.
[0024] To solve the above problems, the inventor found that there are structural defects in the traditional methods for dealing with the coordination of face close-ups and music beats. Through analysis, it is found that the prior art lacks an automated screening mechanism for target faces, and the audio and video processing modules operate independently. From this, the technical concept is generated: to construct a two-layer screening mechanism with face attributes as the core, combine music beat recognition to establish time anchors, and form an audio-visual collaborative editing framework. The specific solution path is: first, establish a basic processing unit through lens segmentation, then implement frame-by-frame screening driven by face attributes, and finally reconstruct the timeline based on the music rhythm to achieve the dynamic adaptation of good-looking shots and music beats.
[0025] Embodiment 1 As Figure 1 shown, a method for automatically processing face-based good-looking video editing in Embodiment 1 of the present invention includes the following steps: Step S100: Obtain the video to be edited and the editing requirements including the target face, and perform lens segmentation on the video to be edited, and cut it into multiple lenses; Step S200: According to the editing requirements including the target face, perform frame extraction analysis on each of the segmented lenses, and analyze whether each frame includes the attribute data of the target face; and extract the frames including the attribute data of the target face; Step S300: Screen the extracted frames including the attribute data of the target face, screen out the frames whose attribute data of the target face meet the editing requirements, and convert them into the lenses of the screened target face; Step S400: Obtain the input specified music as the background music, perform music interval recognition on the specified music, and then perform rhythm point recognition to find the positions where the music makes beats; Step S500: Remove the shots with lines and repeated shots of the filtered target faces, and obtain the filtered shots of the target faces; Step S600: Stitch all the filtered shots of the target faces, stitch the filtered shots of the target faces according to the music beat positions, and edit the stitched video; Step S700: Expand a predetermined duration outward from both sides with the edited video as the center, and select a segment with lines conforming to the current scene and including the target face from the cropped video through a specified large model as the final ending, and generate the final edited video for output.
[0026] After obtaining the video to be edited and the editing requirements including the target face, this application divides the video into shots to form multiple independent shots; performs frame extraction analysis on each shot to extract the frames containing the attribute data of the target face; filters the frames that meet the editing requirements and converts them into target shots; obtains the specified music and identifies its rhythm beat positions; filters the redundant content in the target shots; stitches the shots according to the music beats to form a rough cut video; expands the rough cut video and selects the ending segment that meets the scene, and finally generates a complete video technical solution.
[0027] Among them, shot segmentation refers to the process of dividing a continuous video stream into independent scene units, which can be specifically implemented through a scene change detection algorithm, such as based on the difference of HSV (hue, saturation, value) color histograms or motion vector analysis. This step establishes a structured input for subsequent processing.
[0028] Frame extraction analysis refers to the process of extracting video frames at fixed intervals for feature detection, which can be specifically implemented by combining the face detection module of the OpenCV library with a deep learning model. This process ensures the accurate recognition and attribute extraction of the target face. The target face attribute data includes spatial dimensions, aesthetic scores, and motion features. Specifically, the relative dimensions can be calculated through face key point detection, the beauty degree can be evaluated by the ResNet model, and the motion amplitude can be calculated by the optical flow method. These multi-dimensional data form the quantitative basis for appearance evaluation. Music beat recognition refers to detecting the positions of strong beats in the audio signal, which can be specifically implemented by using the Librosa library to extract spectral features and locating the down-beat points through a peak detection algorithm. This technology supports the accurate alignment of the video rhythm. Shot filtering refers to removing the segments containing dialogues or repeated content, which can be specifically implemented through voice activity detection and image similarity comparison. This step improves the information density of the video content. Video expansion refers to extending the start and end durations of the edited segment, which can be specifically implemented by using the sliding window method to expand the time range. This design provides a buffer interval for the selection of the ending segment.
[0029] Specifically, first, the original video is segmented into independent shot units at scene transition points, and an independent processing channel is established for each shot. The presence state of the target human face is detected by frame-by-frame scanning, and the sequence of valid frames that meet the preset appearance standard is recorded. A timeline reference system is established based on the music rhythm characteristics, and the selected high-appearance shots are arranged and combined according to the rhythm intensity distribution. On the basis of the initial cut video, time expansion is performed, and the semantic relevance of the expanded area is analyzed using a pre-trained language model. The segments that contain both the target human face and conform to the scene context are selected as the end of the video. The whole process forms a closed-loop processing flow, realizing the automatic conversion from the original material to the finished product.
[0030] Compared with the prior art, this solution breaks through the traditional split-track processing mode and establishes a collaborative processing mechanism with the human face attributes as the link and the music rhythm as the backbone. The traditional optical flow method for scene segmentation only considers visual continuity, while this solution performs cross-modal association between the human face aesthetic attributes and the audio rhythm characteristics; the traditional key frame detection relies on global difference analysis, while this solution establishes local feature anchor points through human face attributes; the existing audio beat synchronization technology uses fixed-interval editing, while this solution realizes adaptive rhythm matching based on the dynamically selected high-appearance shots.
[0031] Through the above technical solution, this application effectively solves the problem of automatic editing of videos with the appearance of a specified person, achieving the dual technical effects of precise screening of the target human face and accurate matching of the music rhythm. In the scenario of making short portrait videos, it can automatically retain the shot segments of the best facial expressions of the person, ensure that each picture transition point is synchronized with the strong beats of the music, and at the same time improve the video integrity through the generation of an intelligent ending with semantic association, significantly reducing the time cost and technical requirements of manual editing.
[0032] This application further proposes to perform an overall analysis on the video to be edited, extract all the shots of the video to be edited, and each section between two scene transitions is regarded as a shot.
[0033] Among them, shot segmentation refers to the process of dividing a continuous video stream into multiple independent shots, which can be specifically implemented by inter-frame difference analysis or an image feature clustering algorithm. The scene transition is judged by detecting the mutation points of the picture content. This feature provides independent units for subsequent frame processing through structured segmentation, avoiding interference from cross-shot content in the analysis process.
[0034] Among them, scene transition refers to the conversion node between different shooting scenes or shot languages in the video, which can be specifically implemented by color histogram mutation detection or motion vector mutation recognition methods, and can accurately identify the critical points where the shooting angle, the main body of the picture, or the background environment change significantly. This feature ensures the accuracy of shot division and establishes a reliable time interval basis for subsequent frame extraction based on the target human face.
[0035] Specifically, the system first performs a global scan on the input video and identifies all shot boundary points through a scene change detection algorithm. The sliding window mechanism is used to compare the color distributions of adjacent frames. When the inter-frame color differences exceeding the set threshold continuously appear within the window, it is determined as a scene change. For each continuous video segment between two adjacent scene change points, the system automatically marks it as an independent shot unit, forming a shot set containing complete semantics. This processing method enables subsequent frame extraction analysis to be batch processed in units of shots, improving the operation efficiency.
[0036] Compared with the prior art, traditional video segmentation methods rely on fixed time window cutting or random key frame sampling, which easily leads to incomplete shot segmentation. For example, in sports videos, fixed interval cutting may cut off the continuous actions of athletes, while this solution ensures that each shot contains a complete action segment by dynamically detecting scene change points. Compared with the segmentation scheme based on the optical flow method, which is vulnerable to complex motion interference, this method reduces the computational complexity while ensuring accuracy through color feature mutation detection.
[0037] Through the above technical solution, this application effectively solves the problem of content fragmentation caused by inaccurate shot division in existing video editing technologies. Through accurate scene change detection, shot-level segmentation is achieved, providing complete semantic units for subsequent frame analysis of target faces, avoiding the phenomenon of discontinuous actions or scene jumps caused by incorrect shot division, and significantly improving the quality of basic data for automated editing processing.
[0038] This application further proposes an automated beauty video editing processing method based on faces, which realizes precise screening through multi-dimensional face attribute calculation during the frame extraction analysis process. Specifically, after shot segmentation, frame extraction analysis is performed on each shot. First, it is detected whether each frame image contains a target face, and then for the frames containing the target face, their size data, beauty data, and optical flow value data are calculated respectively to form a complete attribute data set.
[0039] Among them, the size data refers to the relative proportion of the target face in the video frame, which can be specifically calculated by using the bounding box coordinates output by a face detection algorithm based on OpenCV or a deep learning model, and is used to judge the rationality of the face composition in the picture. The face detection algorithm based on OpenCV is a technology that uses the functions provided by the OpenCV library to implement face detection. OpenCV (Open Source Computer Vision Library) is an open-source computer vision and machine learning library, which is widely used in image processing and real-time vision tasks.
[0040] Aesthetics data refers to the quantitative scoring of the facial features of the target face. Specifically, the pre-trained ResNet architecture can be used to extract facial feature vectors and map them into aesthetic score values through the fully connected layer to evaluate the visual appeal of the characters in the picture. Optical flow value data refers to the motion vector statistics of the target face area between consecutive frames. Specifically, the Lucas-Kanade optical flow algorithm can be used to calculate the displacement of feature points between adjacent frames, which is used to analyze the coherence of the character's movements in the shot.
[0041] Specifically, the embodiment of the present application first performs a binary classification judgment on the frame image through a face detection model, and immediately triggers the three-dimensional feature calculation when the target face is detected. When calculating the size of the face, the ratio of the pixel area of the detection frame to the total pixel area of the video frame is used as a quantitative indicator. For example, when the detection frame area ratio is lower than a preset threshold, it can be determined as a long-distance shooting shot. The aesthetic calculation extracts features such as facial features proportions and skin texture through a neural network, and outputs a standardized scoring value for subsequent screening. The optical flow value calculation focuses on the motion trajectory of pixels in the face area, and counts the mean and variance of motion vectors between consecutive frames. For example, when the variance exceeds the critical value, it is determined to be a dynamic expression shot. These three dimensional data jointly construct a quantitative evaluation system for face shots, providing data support for subsequent intelligent screening based on appearance standards.
[0042] Compared with the existing technology, traditional video editing methods mostly use single-dimensional face detection technology, for example, only judging whether the target face exists. This solution innovatively introduces a composite evaluation index, which not only detects the existence of the face, but also establishes a three-dimensional evaluation model through three orthogonal dimensions of size, beauty, and movement. The simple area calculation based on OpenCV in the existing technology cannot accurately reflect the visual expressiveness of the character in the picture, while this solution effectively improves the screening accuracy of the appearance lens by integrating the aesthetic scoring mechanism of deep learning. In addition, the traditional optical flow method is mostly used for overall scene analysis. This solution specifically performs local optical flow calculation on the face area, which can avoid the interference of complex background movement on the dynamic analysis of the character subject.
[0043] Through the above technical solution, this application effectively solves the problems of single screening dimension and poor dynamic adaptability in the processing of appearance videos of traditional editing methods. Through the collaborative analysis of multi-dimensional facial attribute data, high-quality shots that meet the appearance standards and have natural movements can be accurately identified to avoid the phenomenon of misscreening caused by a single detection indicator. Especially when processing video clips containing complex character movements, this solution can accurately distinguish between normal expression changes and interference frames caused by violent movements through the joint analysis of optical flow values and face size, significantly improving the smoothness and viewing experience of video editing, and enhancing the aesthetic feeling of users of video clips.
[0044] This application further proposes steps to screen the frames containing the attribute data of the target face, screen out the frames whose attribute data of the target face meet the editing requirements, and convert them into the shots of the screened target face, which specifically include: screening the frames containing the attribute data of the target face, and comparing the size data, beauty degree data, and optical flow value data of the target face in the frames containing the target face extracted; with the standard appearance size data, standard appearance beauty degree data, and standard appearance optical flow value data of the target face required by the editing requirements respectively; filtering out the frames that do not meet the standard appearance size data, standard appearance beauty degree data, and standard appearance optical flow value data of the target face required by the editing requirements, and screening out the frames whose attribute data of the target face meet the editing requirements; and converting the frames whose attribute data of the target face meet the editing requirements into the shots of the screened target face.
[0045] Among them, the size data of the target face refers to the proportion of the face area in the image, which can be specifically implemented by combining image segmentation algorithms with key point detection techniques. For example, after face area recognition through the Haar cascade classifier or deep learning model in the OpenCV library, the area ratio is calculated. The Haar cascade classifier is a method in the OpenCV library for object detection, especially face detection.
[0046] The beauty degree data of the target face refers to the scored value quantified according to features such as facial symmetry and facial feature proportions, which can be specifically implemented by using a pre-trained aesthetic scoring model. For example, a facial aesthetic evaluation algorithm based on a convolutional neural network extracts features and scores the positional relationships of the eyes, nose, and mouth. The optical flow value data of the target face refers to the motion vector of the face area between adjacent frames, which can be specifically calculated by using the Lucas-Kanade optical flow algorithm or the Farneback dense optical flow method to calculate the degree of motion blur. For example, the motion amplitude is evaluated by tracking the displacement of facial key points in consecutive frames. Among them, the Lucas-Kanade algorithm is a local dense optical flow calculation method used to estimate the motion of objects in an image sequence; it assumes that within a small neighborhood, the motion of the object is consistent and uses the gradient information of the image to deduce the optical flow of each pixel. The Farneback dense optical flow method generates dense optical flow estimates by performing multi-level progressive processing on the image; it uses polynomial fitting to capture the details of the image area containing motion.
[0047] Specifically, the screening process realizes appearance-oriented automated editing by establishing a multi-dimensional standard system. First, three indicators, namely face size, beauty degree, and optical flow value, are extracted from the candidate frames obtained by video frame extraction. Among them, the threshold standard for face size can be set in the range of 10%-30% of the image area, the beauty degree scoring standard can be set above 0.8 points of the model output value, and the optical flow value standard can be limited within the range of 0-5 pixels. For each candidate frame, the three indicators are respectively compared with the preset standard values. For example, when the detected face area ratio of a certain frame is lower than 10%, it is determined that it does not meet the size standard and is excluded; when the beauty degree score of a certain frame is lower than 0.6 points, it is determined that it does not meet the appearance requirements and is filtered. Through the frame-by-frame quantization comparison mechanism, only the high-quality frames that meet all the standard thresholds are retained, and then these screened frames are recombined into continuous shots. For example, the 5th, 7th, and 9th frames that meet the conditions are merged into a 0.5-second shot segment, so as to ensure that the faces presented in the final video meet the preset appearance standards.
[0048] Compared with the prior art, traditional face screening methods usually only rely on single-dimensional features. For example, they only detect the presence of a face or perform simple size filtering, resulting in problems such as blurred faces, unnatural expressions, or dynamic defocus in the selected shots. In contrast, this solution innovatively constructs a multi-dimensional joint screening mechanism. By quantitatively evaluating three key indicators, namely face size, aesthetic quality, and motion stability, it can effectively eliminate low-quality frames. For example, in the scenario of sports event video editing, the prior art may retain the blurred face frames caused by the intense movement of athletes, while this solution can accurately identify such dynamically defocused frames through optical flow value detection for filtering, and at the same time exclude the frames with distorted expressions by combining the beauty degree score, thus significantly improving the material selection quality.
[0049] Through the above technical solution, this application can solve the technical defects of single-dimensional appearance screening and poor adaptability to dynamic scenes in the existing video editing technology, and realize more accurate automated appearance-oriented editing. Especially in the scenario where the shooting object has complex movements, such as dance performances or sports events, through optical flow value detection, dynamically blurred frames can be effectively excluded, and combined with the beauty degree score to ensure the selection of the best facial expressions, so that the generated video achieves an optimized effect in terms of the quality of character presentation, picture clarity, and action coherence, meeting the production needs of short video creators for high-appearance content.
[0050] This application further proposes to obtain the specified input music as the background music, and perform music interval recognition on the specified music, identifying the prelude, verse, pre-chorus, chorus, interlude, bridge, and coda of the specified music; then perform rhythm point recognition to find the down-beat points where the music makes a beat.
[0051] Among them, music interval recognition refers to dividing music into structured sections with different emotional or rhythmic characteristics. Specifically, it can be achieved by combining time-frequency analysis with machine learning models, such as Mel spectrum feature extraction and convolutional neural network classification to distinguish between prelude, verse, coda and other sections. Rhythm point recognition refers to determining the strong beat position by analyzing the periodic beat characteristics of the music signal. Specifically, a signal processing algorithm combined with an autocorrelation function can be used to calculate the fundamental frequency period, or a deep learning model can be used to predict the beat timestamp end-to-end. The down-beat point refers to the strong beat position of the first beat in a music measure. Specifically, it can be determined by a beat tracking algorithm combined with music structure analysis, such as using a dynamic time warping algorithm to align a standard rhythm template.
[0052] Specifically, the input background music is first transformed into a spectrum graph by time-frequency transformation, and the time boundaries of music segments such as the prelude and chorus are identified through the trained classification model. Then, each music segment is rhythmically analyzed, and the beat pulse sequence is extracted by short-time energy calculation combined with peak detection algorithm, and the precise time position of the down-beat point is determined by the phase correction algorithm. In high-energy segments such as choruses, the beat detection accuracy is optimized by increasing the weight coefficient of the spectrum contrast analysis to ensure that the strong beat position matches the emotional climax of the music.
[0053] Compared with the existing technology, the traditional method only divides the music based on a fixed time window or adopts a single beat detection algorithm, which leads to the accumulation of recognition errors of complex music structures. This solution uses a layered processing strategy to first divide the macroscopic music structure and then perform local rhythm analysis, effectively solving the problem of inaccurate detection caused by sudden changes in rhythm features in multi-section music. At the same time, the down-beat point is introduced as the card point benchmark to overcome the defect that the ordinary beat point does not match the emotional expression of music.
[0054] Through the above technical solution, this application realizes the multi-level alignment of music structure and video lens, accurately capturing key rhythm points while retaining the integrity of the music. By locating the down-beat point, it ensures the natural fit between the lens switching action and the strong beat of the music, eliminates the sense of disharmony between the picture switching and the auditory rhythm in traditional editing, and enhances the emotional resonance between the video content and the background music.
[0055] The present application further proposes a method for selecting a high-energy curve of background music, which analyzes the background music of the trailer and the video to be edited, obtains the high-energy curves of the background music of the trailer and the original video to be edited, and defines various rhythm intervals; the method for selecting a high-energy curve of the background music of the trailer includes: analyzing the background music in the time domain or frequency domain to obtain the time corresponding to the chorus part of the background music, thereby obtaining the high-energy curve; obtaining the high-energy curve of the background music according to the sound energy of the background music; and obtaining the high-energy curve based on a model-based method.
[0056] Among them, the time-domain or frequency-domain analysis refers to analyzing the time waveform or spectral characteristics of an audio signal. Specifically, the music structure features can be extracted through the fast Fourier transform or Mel-frequency cepstral coefficients to locate the chorus time interval in a music passage. The sound energy analysis refers to calculating the amplitude intensity distribution of the audio signal in the time dimension. Specifically, the short-time energy calculation method can be used to identify the high-energy interval corresponding to the climax part of the music. The model-based method refers to training an audio feature recognition model using a deep learning framework. Specifically, a convolutional neural network or a recurrent neural network can be used to learn the music energy distribution pattern, so as to predict the shape of the high-energy curve.
[0057] Specifically, during the music interval recognition stage, the high-energy curve data is constructed synchronously. The music structure features are extracted through time-domain analysis to determine the chorus position, and the basic high-energy curve is obtained by combining the frequency-domain energy distribution calculation. Then, the curve shape is optimized and corrected through a pre-trained model. The high-energy curve of the original video is generated by analyzing the shot transition frequency and the picture motion intensity, and is matched with the high-energy curve of the background music in the rhythm interval to form a correspondence table between the music energy peak and the video dynamic high point, providing a time alignment benchmark for subsequent shot splicing.
[0058] Compared with the prior art, the traditional method only relies on a single beat detection algorithm to process the audio signal, and there are deviations in the rhythm recognition of non-periodic sound components and complex music structures. This solution can accurately capture the significant music energy fluctuations perceived by the human ear by integrating a multi-dimensional high-energy curve generation mechanism and introducing a machine learning model correction mechanism in the music structure analysis stage. At the same time, it combines the video dynamic features to generate a matching rhythm interval division rule, effectively solving the problem of the disconnection between the audio-visual features.
[0059] Through the above technical solution, this application realizes the deep coupling of the music rhythm features and the video dynamic features, ensuring that the shot transition points simultaneously meet the requirements of music positioning and the need for picture coherence. Specifically, the high-energy curve matching algorithm can automatically align the music climax section with the strongly moving video segments, eliminating the phenomenon of misalignment between the picture transition and the music beat in traditional editing, and improving the rhythm coordination and viewing fluency of video editing.
[0060] This application further proposes steps of splicing all the filtered target face shots, splicing the filtered target face shots according to the music beat positions, and editing the spliced video, which include: obtaining the filtered target face shots, organizing the filtered target face shots, and classifying them according to chronological order or emotional expression; obtaining the identified music beat positions, marking the down-beats of strong beats and the music beat positions of the climax parts according to the music rhythm; splicing the organized target face shots on the appropriate beats according to the previously identified music beat positions to ensure that each shot transition is synchronized with the music rhythm; editing the spliced video, automatically adjusting the time of shot transitions, special effects, and transition effects to ensure the smooth coordination of the entire video.
[0061] Among them, the filtered target face shots refer to the sequence of target face pictures obtained by removing lines and repeated shots, which can be specifically implemented by using speech recognition and image similarity algorithms, and are used to retain high-quality and diverse segments without interference. The music beat positions refer to the nodes with significant rhythm changes in the background music, which can be specifically implemented by using Mel Frequency Cepstral Coefficients and peak detection algorithms, and are used to determine the precise time points of shot transitions. The classification according to chronological order or emotional expression means grouping the shots according to the time line or emotional tags, which can be specifically implemented by using timestamp matching and emotion recognition models, and is used to maintain the logical coherence of the video. The synchronization of shot transitions means aligning the picture transitions with the music beats, which can be specifically implemented by using a time axis mapping algorithm, and is used to eliminate the problem of out-of-sync audio and video. Automatically adjusting the transition effects means adapting the transition special effects according to the rhythm intensity, which can be specifically implemented by using a preset special effect template and beat intensity matching algorithm, and is used to enhance the audio-visual consistency.
[0062] Specifically, the filtered target face shots are first classified and organized through timeline analysis or emotion recognition models to form a shot sequence with internal logic. The music beat positions are extracted for strong beats and climax intervals through audio signal processing technology to generate rhythm marking points accurate to the millisecond level. During the shot splicing process, the dynamic time warping algorithm is used to align the starting points of each shot with the music beats, and at the same time, the fade-in / fade-out or fast-switching transition methods are automatically selected according to the beat intensity. For example, a fast switch of 0.1 second is used to match the dense drum beats in the climax part of the chorus, and a 1-second fade transition is used to match the soothing melody in the prelude. After splicing, the video editing engine automatically optimizes the shot connection to eliminate black frames or audio-visual delay phenomena.
[0063] Compared with the prior art, traditional methods only simply associate shot transitions with musical beats, ignoring the matching relationship between shot content and musical emotions, which easily leads to incongruous phenomena such as slow-paced music being paired with action segments. This solution uses a dual classification mechanism to superimpose an emotional matching dimension on the basis of time sorting. For example, a close-up shot of a smile is corresponded to a bright melody passage, and a profile shot of a face is corresponded to a solo string passage, achieving dual synchronization of content and emotion. In addition, the prior art fixedly uses a single transition effect, while this solution dynamically adjusts according to the beat intensity. For example, a flash transition is used for strong beats and a smooth dissolve is used for weak beats, effectively enhancing the rhythm expressiveness.
[0064] Through the above technical solution, this application solves the problem of serious disconnection between audio and video in traditional editing tools, enabling high-quality shots of the target person to precisely match the emotional ups and downs of the music. For example, in the editing of dance videos, the apex of the rotating action shot automatically aligns with the downbeat rhythm, and the duration of the close-up shot of the expression is synchronized with the duration of the chord, thereby generating short videos with professional-level rhythm. At the same time, real-time processing on the mobile side is achieved through a lightweight algorithm, and multi-dimensional feature matching can be completed without relying on a high-performance GPU, significantly improving the consistency of editing efficiency and finished product quality.
[0065] The following further elaborates on the present invention through specific application embodiments: Embodiment 2 As Figure 2 shown, an automated high-quality video editing and processing method based on human faces provided by this specific application embodiment 2 includes the following steps: S40. Audio medium analysis step: Input an audio medium, and perform audio beat (rhythm) recognition and music interval recognition on the input audio medium such as a specified piece of music; For example, perform audio beat analysis (intro, verse, pre-chorus, chorus, interlude, bridge, outro) on the specified piece of music, and find the positions (down-beat points) where the music can be used for beat matching; The audio medium analysis can be completed in one minute (1min); Among them, intro: The beginning part of the audio, usually used to introduce the audio theme and set the atmosphere.
[0066] Verse: The main narrative part of the audio, usually containing the story or emotional core of the audio, and usually repeated multiple times.
[0067] Pre-chorus: The part connecting the verse and the chorus, used to increase the tension and guide the listener into the chorus.
[0068] Chorus: The most core and memorable part of the audio, with the lyrics repeated, conveying the main theme or emotion.
[0069] Interlude: The instrumental part between the chorus or the main verse, providing a break and adding layers to the music.
[0070] Bridge: A contrasting section that usually appears once in a song, providing a different melody or emotion to highlight the change in the song.
[0071] Outro: The ending part of a song, used to conclude the whole song, usually reviewing or reaffirming the theme.
[0072] S41. Perform video medium analysis, input the video medium, and segment the shots of the input video medium.
[0073] For example, input the video to be edited (specified film), segment the shots of the input specified film. Shot segmentation means analyzing the whole film, extracting all the shots of the film, and dividing them into multiple shots such as shot1, shot2... shotN. The section between every two scene switches is regarded as a shot. The shot segmentation process can be completed in 10 minutes. S42. Perform frame extraction analysis on the video medium, extract each frame of the video for portrait recognition of the target face, and convert the frame recognition data into shot recognition data. For example, perform frame extraction and analysis on the film, extract each frame of the video to calculate the required data, including the target face, the size of the target face, the beauty degree, and the optical flow value.
[0074] S43. Perform subtitle OCR (media assets) on the input video medium. Performing subtitle OCR (Optical Character Recognition) on the input video medium (the video to be edited) means using OCR technology to recognize the shots with subtitles in the video.
[0075] The following are the steps for assembling the mixed-cut video: S44. Select the shots containing the specified actor and proceed to S45.
[0076] S45. Eliminate the shots with lines according to the subtitle OCR recognition and proceed to S46.
[0077] S46. Eliminate duplicate shots based on image similarity; and proceed to S47; S47. Assemble the video according to the music beats and proceed to S48; S48. Video assembly, adding audio and video effects. Centering on the last shot segment after splicing, expand a certain duration to both sides, and use the large model to select a segment with better lines and the target as the main body from it as the final ending.
[0078] In this specific embodiment, the target face shots after removing the shots with lines and repeated shots are first classified and sorted through timeline analysis or an emotion recognition model to form a shot sequence with internal logic. The music beat positions are extracted for strong beats and climax intervals through audio signal processing technology to generate rhythm marker points accurate to the millisecond level. During the shot splicing process, the dynamic time warping algorithm is used to align the starting points of each shot with the music beats, and at the same time, the fade-in / fade-out or fast-switching transition method is automatically selected according to the beat intensity. For example, a 0.1-second fast switch is used to match the dense drum beats in the climax part of the chorus, and a 1-second gradual transition is used to match the soothing melody in the intro part. After splicing, the video editing engine automatically optimizes the shot connection to eliminate black frames or audio-visual delay phenomena.
[0079] Compared with the prior art, traditional methods only simply associate shot transitions with music beats, ignoring the matching relationship between shot content and music emotions, which easily leads to inconsistent phenomena such as slow-paced music for action segments. This solution uses a dual classification mechanism to superimpose the emotional matching dimension on the basis of time sorting. For example, a close-up shot of a smile is corresponding to a bright melody passage, and a profile shot of the side face is corresponding to a solo string passage, achieving double synchronization of content and emotion. In addition, the prior art fixedly uses a single transition special effect, while this solution dynamically adjusts according to the beat intensity. For example, a flash transition is used for strong beats and a smooth dissolve is used for weak beats, effectively enhancing the rhythm expressiveness.
[0080] Among them, in the embodiment of the present invention, the climax interval is obtained by extracting the high-energy curve of the selected audio; Among them, the method for selecting the high-energy curve of BGM (background music): The high-energy curves of the trailer BGM and the film can be analyzed respectively to obtain the high-energy curves of the trailer BGM and the original film, and various rhythm intervals are delimited; And the method for selecting the high-energy curve of the trailer BGM (background music) is as follows: 1). Analyze the BGM (background music) in the time domain or frequency domain to obtain the time corresponding to the chorus part of the BGM (background music), and thus obtain the high-energy curve (in fact, the chorus part of a piece of music is basically the high-energy part); 2). Obtain the BGM high-energy curve according to the sound energy (it may be okay because the volume of the climax part of a piece of music is higher than that of the non-climax part); 3). Based on a model method, such as deepchorus (this task is actually to obtain the chorus part of the BGM) to obtain the high-energy curve.
[0081] Exemplary device As Figure 3 shown, the embodiment of the present invention provides a face-based automated beauty video editing and processing device, which includes: The shot segmentation module 310 is used to obtain the video to be edited and the editing requirements including the target face, segment the video to be edited into multiple shots; The target face frame extraction module 320 is used to perform frame extraction analysis on each segmented shot according to the editing requirements including the target face, analyze whether each frame includes the attribute data of the target face; and extract the frames including the attribute data of the target face; The target face frame screening module 330 is used to screen the frames including the attribute data of the target face, screen out the frames whose attribute data of the target face meet the editing requirements, and convert them into the shots of the filtered target face; The music beat recognition module 340 is used to obtain the input specified music as the background music, identify the music intervals of the specified music, and then perform rhythm point recognition to find the positions where the music beats; The target face shot filtering module 350 is used to remove the shots with lines and repeated shots from the filtered shots of the target face to obtain the filtered shots of the target face; The target face shot splicing module 360 is used to splice all the filtered shots of the target face, splice the filtered shots of the target face according to the music beat positions, and edit the spliced video; The video extension and output module 370 is used to expand a predetermined duration outward from both sides with the edited video as the center, and select a segment of the line that conforms to the current scene and includes the target face from the cropped video through a specified large model as the final ending, and generate the final edited video output 370, as described above.
[0082] Based on the above embodiments, the present invention also provides an intelligent terminal, and its principle block diagram can be as Figure 4 shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an automated beauty video editing processing method based on the face. The database of the intelligent terminal is used to store the automated beauty video editing processing program based on the face.
[0083] Those skilled in the art can understand, Figure 4The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0084] In one embodiment, an intelligent terminal is provided, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain the video to be clipped and the clipping requirements including the target face, perform shot segmentation on the video to be clipped, and divide it into multiple shots; according to the clipping requirements including the target face, perform frame extraction analysis on each of the segmented shots to analyze whether each frame includes the attribute data of the target face; and extract the frames including the attribute data of the target face; screen the extracted frames including the attribute data of the target face, screen out the frames whose attribute data of the target face meet the clipping requirements, and convert them into the shots of the filtered target face; obtain the input specified music as the background music, perform music interval recognition on the specified music, and then perform rhythm point recognition to find the positions where the music makes beats; remove the shots with lines and repeated shots from the filtered shots of the target face to obtain the filtered shots of the target face; splice all the filtered shots of the target face, splice the filtered shots of the target face according to the music beat positions, and clip the spliced video; expand outward by a predetermined duration with the clipped video as the center, and select a segment of lines that conforms to the current scene and includes the target face from the cropped video through a specified large model as the final ending to generate the final clipped video for output.
[0085] Further, the step of performing shot segmentation on the video to be clipped and dividing it into multiple shots includes: performing an overall analysis on the video to be clipped, and extracting all the shots of the video to be clipped, where the section between every two scene switches is regarded as one shot.
[0086] Further, according to the editing requirements including the target face, frame extraction analysis is performed on each of the segmented shots to analyze whether each frame includes the attribute data of the target face; and the steps of extracting the frames including the attribute data of the target face include: wherein, the attribute data of the target face are: the target face, as well as the size data, aesthetics data, and optical flow value data of the target face; according to the editing requirements including the target face, frame extraction analysis is performed on each of the segmented shots to analyze whether each frame includes the target face, and the frames including the target face are extracted; for the extracted frames including the target face, the size data, aesthetics data, and optical flow value data of the target face are analyzed and calculated to obtain the frames including the attribute data of the target face.
[0087] Further, the steps of screening the frames including the attribute data of the target face, screening out the frames whose attribute data of the target face meet the editing requirements, and converting them into the shots of the screened target face include: screening the frames including the attribute data of the target face, and taking the size data, aesthetics data, and optical flow value data of the target face in the extracted frames including the target face; comparing them with the standard appearance size data, standard appearance aesthetics data, and standard appearance optical flow value data of the target face required by the editing requirements respectively; filtering out the frames that do not meet the standard appearance size data, standard appearance aesthetics data, and standard appearance optical flow value data of the target face required by the editing requirements, and screening out the frames whose attribute data of the target face meet the editing requirements; and converting the screened frames whose attribute data of the target face meet the editing requirements into the shots of the screened target face.
[0088] Further, the steps of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the position where the music makes a beat include: obtaining the input specified music as the background music, performing music interval recognition on the specified music, and recognizing the prelude, verse, guide song, chorus, interlude, bridge, and coda of the specified music; then performing rhythm point recognition to find the down-beat point where the music makes a beat.
[0089] Further, the steps of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the position where the music makes a beat also include: Method for selecting the high-energy curve of the background music: Analyze the trailer background music and the video to be edited, respectively obtain the high-energy curves of the trailer background music and the original video to be edited, and delimit various rhythm intervals; The method for selecting the high-energy curve of the trailer background music includes: performing time-domain or frequency-domain analysis on the background music to obtain the time corresponding to the chorus part of the background music, so as to obtain the high-energy curve; obtaining the high-energy curve of the background music according to the sound energy of the background music; and obtaining the high-energy curve based on the model method.
[0090] Further, the steps of splicing all the filtered target face shots, splicing the filtered target face shots according to the music beat positions, and editing the spliced video include: obtaining the filtered target face shots, organizing the filtered target face shots, and classifying them according to chronological order or emotional expression; obtaining the identified music beat positions, marking the down-beats of strong beats and the music beat positions of climax parts according to the music rhythm; splicing the organized target face shots on appropriate beats according to the previously identified music beat positions to ensure that each shot transition is synchronized with the music rhythm; editing the spliced video, automatically adjusting the time, special effects, and transition effects of shot transitions to ensure that the whole video is coordinated and smooth, as specifically described above.
[0091] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0092] In summary, the present invention provides a method, apparatus, intelligent terminal and storage medium for automated video editing and processing based on human faces. Through techniques such as lens segmentation, target human face attribute analysis, music beat synchronization and intelligent expansion, it realizes video editing with a high degree of matching between human face beauty and music rhythm, solves the problems of serious audio-visual fragmentation and insufficient dynamic adaptability in traditional methods, and has the effects of improving editing smoothness and enhancing video expressiveness; it can automatically edit and generate high-quality short videos with target human faces, meet the needs of creators and the market, improve video editing efficiency and quality; the present invention can simultaneously improve the accuracy of video editing; the generated video images will correspond to the high points of the audio, achieving the purpose of attracting more users to watch.
Claims
1. An automated face-based video clip processing method for facial attractiveness, characterized in that Including: Obtain the video to be clipped and the clipping requirements including the target face, perform shot segmentation on the video to be clipped, and segment it into multiple shots; According to the clipping requirements including the target face, perform frame extraction analysis on each segmented shot to analyze whether each frame includes the attribute data of the target face; and extract the frames including the attribute data of the target face; Screen the extracted frames including the attribute data of the target face, screen out the frames whose attribute data of the target face meet the clipping requirements, and convert them into the shots of the target face after screening; Obtain the input specified music as the background music, perform music interval recognition on the specified music, and then perform rhythm point recognition to find the positions where the music makes beats; Remove the shots with lines and duplicates from the screened shots of the target face to obtain the filtered shots of the target face; Stitch all the filtered shots of the target face, stitch the filtered shots of the target face according to the music beat positions, and edit the stitched video; Taking the edited video as the center, expand a predetermined duration to both sides, and select a segment with lines conforming to the current scene and including the target face from the cropped video through a specified large model as the final ending to generate and output the final clipped video.
2. The automated beauty video editing and processing method based on human face according to claim 1, wherein, The step of performing shot segmentation on the video to be clipped and segmenting it into multiple shots includes: Perform an overall analysis on the video to be clipped, and extract all the shots of the video to be clipped. Among them, the section between every two scene switches is regarded as a shot.
3. The automated face-based video clip processing method for appearance value according to claim 1, characterized in that According to the clipping requirements including the target face, perform frame extraction analysis on each segmented shot to analyze whether each frame includes the attribute data of the target face; And the step of extracting the frames including the attribute data of the target face Including: Among them, the attribute data of the target face are: the target face, as well as the size data, beauty degree data, and optical flow value data of the target face; According to the clipping requirements including the target face, perform frame extraction analysis on each segmented shot to analyze whether each frame includes the target face, and extract the frames including the target face; For the extracted frames including the target face, analyze and calculate the size data, beauty degree data, and optical flow value data of the target face to obtain the frames including the attribute data of the target face.
4. The automated face-based beauty video editing and processing method according to claim 1, characterized in that The step of screening the extracted frames including the attribute data of the target face, screening out the frames whose attribute data of the target face meet the clipping requirements, and converting them into the shots of the target face after screening includes: Screen the extracted frames including the attribute data of the target face, and compare the size data, beauty degree data, and optical flow value data of the target face in the extracted frames including the target face with the standard appearance size data, standard appearance beauty degree data, and standard appearance optical flow value data required by the clipping requirements respectively; Frames that do not meet the standard appearance size data, standard appearance beauty data, and standard optical flow value data of the target face required by the described editing requirements are filtered out, and frames whose attribute data of the target face meet the described editing requirements are selected; And the frames whose attribute data of the target face meet the described editing requirements, which are selected, are converted into shots of the filtered target face.
5. The automated face-based video clip processing method for appearance value according to claim 1, characterized in that, The steps of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the positions where the music makes beats include: Obtain the input specified music as the background music, and perform music interval recognition on the specified music to recognize the prelude, verse, guide song, chorus, interlude, bridge, and coda of the specified music; Then perform rhythm point recognition to find the down-beat points, the positions where the music makes beats.
6. The automated face-based video clip processing method for appearance value according to claim 1, wherein The steps of obtaining the input specified music as the background music, performing music interval recognition on the specified music, and then performing rhythm point recognition to find the positions where the music makes beats further include: Method for selecting the high-energy curve of the background music: Analyze the trailer background music and the video to be edited, respectively obtain the high-energy curves of the trailer background music and the original video to be edited, and delimit various rhythm intervals; The method for selecting the high-energy curve of the trailer background music includes: Perform time-domain or frequency-domain analysis on the background music to obtain the time corresponding to the chorus part of the background music, and thus obtain the high-energy curve; Obtain the high-energy curve of the background music according to the sound energy of the background music; Based on the model method, obtain the high-energy curve.
7. The automated facial attractiveness video editing and processing method according to claim 1, wherein The steps of splicing all the filtered shots of the target face, splicing the filtered shots of the target face according to the music beat positions, and editing the spliced video include: Obtain the filtered shots of the target face, and organize the filtered shots of the target face, classifying them according to time sequence or emotional expression; Obtain the found music beat positions, and mark the strong beats down-beat and the music beats at the climax part according to the music rhythm; According to the previously recognized music beat positions, splice the organized shots of the target face at the appropriate beats to ensure that each shot transition is synchronized with the music rhythm; Edit the spliced video, automatically adjust the time, special effects, and transition effects of the shot transitions to ensure that the entire video is coordinated and smooth.
8. An automated facial appearance video editing and processing device based on human faces, characterized in that, The device includes: A shot segmentation module, configured to obtain the video to be edited and obtain the editing requirements including the target face, and segment the video to be edited into multiple shots; A target face frame extraction module, configured to perform frame extraction analysis on each segmented shot according to the editing requirements including the target face, analyze whether each frame includes the attribute data of the target face; and extract the frames including the attribute data of the target face; A target face frame screening module, configured to screen the extracted frames including the attribute data of the target face, screen out the frames whose attribute data of the target face meet the editing requirements, and convert them into the shots of the filtered target face; A music beat recognition module, which is used to obtain the specified input music as the background music, identify the music intervals of the specified music, then perform rhythm point recognition, and find the positions where the music makes beats; A target face shot filtering module, which is used to filter out the shots with lines and repeated shots of the target face after screening, and obtain the filtered shots of the target face; A target face shot splicing module, which is used to splice all the filtered shots of the target face, splice the filtered shots of the target face according to the music beat positions, and edit the spliced video; A video extension and output module, which is used to expand a predetermined duration outward from both sides with the edited video as the center, and select a segment with lines that match the current scene and include the target face from the cropped video through a specified large model as the final ending, and generate and output the final edited video.
9. An intelligent terminal, characterized in that, It includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Video editing method based on deep learning, related equipment and storage medium
CN113709384A
Video material screening method and device, equipment and medium
CN114780795A
Video data processing method and device, electronic equipment and storage medium
CN117156078A
Video generation method and system based on music
CN117412094A
Video retrieval method and device
CN119597969A
Cited By
Dance video generation method, device and equipment and readable storage medium
CN121037644A