Film and television video editing processing method and device and computer storage medium
By extracting the dynamic features of video materials and building a material relationship map, generating timing lines and automatically splicing video materials, the abrupt problem of material switching in video clips is solved, and an efficient and automated video editing process is achieved.
Patent Information
- Application Number
- CN202510511090.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the video editing process, the directly spliced video materials have a sudden problem of switching between materials, resulting in low video editing efficiency.
By extracting the dynamic features of video materials, a material relationship map is constructed, and timing lines are generated based on the map, the video materials are automatically spliced and the transition method is customized.
It realizes the entire process automation of video editing, quickly generates the optimal editing sequence, and significantly improves the efficiency and visual coherence of video editing.
Smart Images

Figure CN120128769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method, device and computer storage medium for editing and processing film and television videos. Background Art
[0002] With the rapid development of Internet technology, videos or images can be edited through applications in electronic devices to obtain videos with rich visual effects. In the video editing scenario, users want to splice multiple video materials to obtain a continuous short film. However, the directly spliced videos often have the problem of abrupt switching between materials. Therefore, at present, the time sequence layout of each video material is often manually performed to reduce the abruptness during material switching. Using this method, the video editing efficiency is low.
[0003] Therefore, it is necessary to provide a method, device and computer storage medium for editing and processing film and television videos to solve the above technical problems. Summary of the Invention
[0004] The present invention overcomes the deficiencies of the prior art and provides a method, device and computer storage medium for editing and processing film and television videos.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a method for editing and processing film and television videos, including the following steps:
[0006] Obtain a plurality of video materials;
[0007] Extract the dynamic effect features of each of the video materials;
[0008] Construct a material relationship graph based on the similarity of the dynamic effect features of each of the video materials;
[0009] Generate a time sequence line according to the degree of association between each of the video materials on the material relationship graph;
[0010] Splice the corresponding video materials in sequence according to the time sequence line, and customize the transition method between adjacent video materials to obtain an edited video.
[0011] In a preferred embodiment of the present invention, the method for extracting the dynamic effect features includes the following steps:
[0012] Preprocess the video material, including grayscale conversion, noise reduction and smoothing, and size adjustment;
[0013] Use the Lucas-Kanade algorithm to calculate the optical flow field of the video material;
[0014] Extract the dynamic effect features from the optical flow field.
[0015] In a preferred embodiment of the present invention, the dynamic effect features include: the movement trajectory, speed, direction, and position of the main body in the video frame;
[0016] Movement trajectory: Based on the prev_pts and curr_pts of the Lucas-Kanade algorithm, track the feature points and save the coordinate sequence of each point;
[0017] Speed: where u and v are the components of the optical flow vector;
[0018] Direction:
[0019] Position: Select the coordinate sequences of several points in the movement trajectory, including at least the coordinate sequences at both ends of the movement trajectory.
[0020] In a preferred embodiment of the present invention, the calculation of the similarity between the dynamic effect features of the video material includes the following steps:
[0021] Convert the dynamic effect feature data of the video material into a numerical vector for calculating similarity, and calculate the similarity respectively:
[0022] Feature Vector = {trajectory similarity, speed similarity, direction similarity, position similarity}, specifically:
[0023] Trajectory similarity: Use dynamic time warping (DTW) to calculate the matching distance between two trajectories: where T A , T B are the coordinate sequences of the two trajectories, N is the length of the trajectory, and i is the i-th point on the movement trajectory;
[0024] Speed similarity: Calculate the Euclidean distance of the speed sequence:
[0025] where v A (t), v B (t) are the speed values of the two materials at the t-th frame;
[0026] Direction similarity: Calculate the cosine similarity of the direction vectors: DirSim(θ A , θ B ) = cos(θ A - θ B ), where θ A , θ B are the average direction angles of the two video materials;
[0027] Position similarity: Compare the coordinate distances between the end frame and the start frame:
[0028]
[0029] Furthermore, fuse multiple feature similarities through weighted summation: Similarity(A,B) = w 轨迹 ·
[0030] DTW(T A ,T B ) + w 速度 ·SpeedSim(S A ,S B ) + w 方向 ·DirSim(θ A ,θ B ) + w 位置 ·
[0031] DirSim(θ A ,θ B ), where w 轨迹 , w 速度 , w 方向 , w 位置 are preset weights;
[0032] Normalize the similarity value to the interval [0, 1].
[0033] In a preferred embodiment of the present invention, the method for constructing the material relationship graph includes:
[0034] Each of the video materials is correspondingly set with a node to store the dynamic effect feature vector: Node(i) = {ID, trajectory, speed sequence, direction sequence, position coordinates};
[0035] Traverse all pairs of the video materials, and use the similarity value as the edge: Edge(i,j) = {Node i , Node j , Similarity(i,j)}, to construct an initial relationship graph;
[0036] Set a similarity threshold, and remove the weak association edges with similarity lower than the threshold on the initial material relationship graph to obtain the material relationship graph.
[0037] In a preferred embodiment of the present invention, the generation of the time series line includes the following steps:
[0038] Randomly select several high centrality nodes as candidate starting points to generate an initial path pool;
[0039] Adopt the weighted longest path algorithm. For the end node of each path, traverse its unvisited neighbor nodes, expand the path in descending order of edge weights, calculate the cumulative quality function Q of the path after each expansion, and only retain the M paths with the highest Q value;
[0040] Cumulative mass function Q: where the weight coefficients are α = 0.5, β = 0.3, and γ = 0.2;
[0041] Stop expanding when the path covers all nodes or the relevance of the remaining nodes is lower than the threshold, and generate several time series lines including the order of the video materials.
[0042] In a preferred embodiment of the present invention, according to the difference in the dynamic effect characteristics of adjacent video materials, the transition type is dynamically selected, specifically:
[0043] If the motion trajectory directions of adjacent video materials are continuous, use optical flow interpolation to generate a motion blur transition;
[0044] If the main body positions of adjacent video materials are complementary, generate a dynamic mask according to the positions;
[0045] If the motion directions of adjacent video materials are opposite, insert a direction reversal special effect;
[0046] If the speed difference between adjacent video materials is significant, insert a speed fade animation.
[0047] In a preferred embodiment of the present invention, the determination of the start frame and the end frame of each video material includes the following steps:
[0048] For each pair of adjacent video materials {A i , A i+1}, load the dynamic effect feature data of both;
[0049] Through the key frame matching algorithm, screen out the candidate key frames strongly related to the motion trajectory from each video material, and determine the end frame of A i with the most matching motion characteristics and the start frame of A i+1 between the candidate end frame of A i and the candidate start frame of A i+1 ;
[0050] According to the determined start frame and end frame of each video material, clip the corresponding video material.
[0051] An electronic device, comprising:
[0052] A processor;
[0053] A memory; and
[0054] Computer program instructions stored in the memory, which, when run by the processor, cause the processor to execute the video and film clip processing method according to any one of claims 1-8.
[0055] A computer storage medium includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the video and film video clip processing method described in any one of the above.
[0056] The present invention solves the defects existing in the background art, and the present invention has the following beneficial effects:
[0057] (1) The present invention provides a video and film video clip processing method. Through the automatic extraction and analysis of dynamic effect features (including motion trajectory, speed, direction, and position), and in combination with a material relationship graph and a weighted path algorithm, it realizes the full-process automatic operation from material selection to the output of the final edited product. This method can not only quickly generate the optimal editing sequence, greatly shortening the time required to process a large number of materials and significantly improving the overall efficiency of video editing, but also maintain a high degree of professionalism and accuracy during the processing. Moreover, by dynamically matching the dynamic effect features of adjacent materials and intelligently customizing the transition effects, it ensures that the picture connection is natural and smooth, significantly enhancing the visual coherence and artistic expressiveness of the video.
[0058] (2) The material relationship graph constructed by the present invention visualizes the similarity of the dynamic effect features of video materials as an association network through a graph structure of nodes and edges. The nodes represent individual video materials, the edges represent the intensity of similarity between materials, and the weight of the edges quantifies their degree of association (such as the matching degree of trajectory, speed, and direction). It can replace manual frame-by-frame matching, especially suitable for the efficient processing of hundreds of materials. And based on the edge weights of the graph, the optimal timing line is generated through a path search algorithm to ensure the coherence of the spliced materials in terms of motion trajectory, speed, and direction.
[0059] (3) When the present invention splices video materials according to the generated timing line, through a key frame matching algorithm, candidate key frames strongly related to the motion trajectory are screened out from each video material, and position matching is performed. The candidate end frame and candidate start frame with the closest positions are selected as the end frame of the previous video material and the start frame of the next material, so that the main body motion trajectories are as connected as possible after the material switch, to ensure the coherence of the video after the editing process. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;
[0061] Figure 1 It is a flowchart of the video and film video clip processing method of the preferred embodiment of the present invention;
[0062] Figure 2 It is a flowchart of the method for extracting the dynamic effect features of the preferred embodiment of the present invention;
[0063] Figure 3 It is a flowchart of the method for generating the timing line of the preferred embodiment of the present invention;
[0064] Figure 4 It is a block diagram of the electronic device of the preferred embodiment of the present invention. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0066] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0067] Application overview
[0068] In the era of self-media, the number of various short videos on the Internet has increased exponentially. Among them, there is a type of video that combines multiple videos with similar main body actions of various pictures through editing means to produce a humorous and visually rich video. Currently, basically, users make videos by dragging materials to the timeline and manually adjusting the order and transition effects. However, when processing hundreds of materials, manual layout is time-consuming and the video editing efficiency is low.
[0069] To solve the above problems, the present invention extracts the dynamic effect features of video materials (such as information on the movement trajectory, speed, direction, position, etc. of the main body) through computer vision technology, and constructs a material relationship graph based on these features, which can effectively analyze the relevance between materials. The editing timing line generated based on the degree of association between materials can guide the splicing order of video materials, and can synchronously process the splicing of multiple materials, greatly improving the editing efficiency of this type of video. At the same time, the naturalness of the switching is optimized through the dynamically generated transition effects, thereby enhancing the overall visual effect of the video.
[0070] After introducing the basic principle of the present invention, the various non-limiting embodiments of the present invention will be specifically introduced below with reference to the accompanying drawings.
[0071] Exemplary method
[0072] Figure 1The figure shows a flowchart of a method for processing film and television video clips according to an embodiment of the present invention.
[0073] As Figure 1 shown, the method for processing film and television video clips according to an embodiment of the present application includes: S1, obtaining a plurality of video materials; S2, extracting the dynamic effect features of each video material; S3, constructing a material relationship graph based on the similarity of the dynamic effect features of each video material; S4, generating a time sequence line according to the degree of association between each video material on the material relationship graph; S5, sequentially splicing the corresponding video materials according to the time sequence line, and customizing the transition method between every two video materials to obtain a clipped video.
[0074] The present invention significantly improves the efficiency and quality of video clipping through computer vision-driven automated clipping technology, and solves the pain points of traditional manual clipping, such as long time consumption, dependence on experience, and insufficient visual coherence.
[0075] Next, each step will be described in detail.
[0076] Step S1, obtaining a plurality of video materials.
[0077] The obtained video materials can be film and television contents in different time periods of the same film and television work, can be a section of film and television content in different film and television works, or can also be a plurality of videos shot by users.
[0078] It should be noted that the obtained video materials should be in the same format, such as mp4 or AVI.
[0079] Step S2, extracting the dynamic effect features of each video material.
[0080] That is to say, computer vision technology is used to analyze each video material to extract the dynamic effect features of each video material. Among them, the dynamic effect features are key parameters extracted from the video material to describe the dynamic changes of the picture, and are used to analyze the motion law, main body behavior and scene changes of the video content, and are the core basis for realizing the automated clipping, intelligent transition and special effect generation of the present invention.
[0081] The dynamic effect features include: information such as the motion trajectory, speed, direction, and position of the main body in the video picture.
[0082] Motion trajectory: The path that the main body (such as a person or an object) in the picture moves over time. For example, when a person waves a wooden stick, the coordinate sequence of the tip of the wooden stick moving from the upper left corner to the lower right corner of the picture. Speed: The rate of movement of the main body, expressed in pixels / frame or physical units (such as m / s), to judge the speed of movement and is used for dynamic matching during transitions.
[0083] Direction: The direction vector of the main body's movement (such as horizontal, vertical, oblique), calculated by computing the tangent direction of the movement trajectory or the direction component of the optical flow field.
[0084] Position: The coordinates of the main body in each frame, analyzing the movement range of the main body in the picture, used for matte animation or spatial connection.
[0085] Figure 2 The flowchart of the method for extracting dynamic effect features according to an embodiment of the present invention is shown. The dynamic effect features of the video material are extracted by the optical flow method, specifically including the following steps:
[0086] Step S21: Preprocess the video material, including grayscale conversion, noise reduction and smoothing, and size adjustment. The size adjustment is to scale the large-resolution video to balance accuracy and computational efficiency.
[0087] Step S22: Calculate the optical flow field using the Lucas-Kanade algorithm.
[0088] Step S23: Extract the dynamic effect features from the optical flow field calculated in step S22, including:
[0089] Movement trajectory: Based on prev_pts and curr_pts of the Lucas-Kanade algorithm, track feature points (such as corner points) and save the coordinate sequence of each point.
[0090] Speed calculation: (Pixels / frame), where u and v are the components of the optical flow vector.
[0091] Direction calculation: where the angle θ deg Range:
[0092] 0°: Along the positive x-axis (right);
[0093] 90°: Along the positive y-axis (down);
[0094] 180°: Along the negative x-axis (left);
[0095] 270°: Along the negative y-axis (up).
[0096] Position: Select the coordinate sequences of several points in the movement trajectory, including at least the coordinate sequences at both ends of the movement trajectory.
[0097] Furthermore, the position can also select the points with significant speed changes or sudden direction changes in the movement trajectory, which is convenient for processing the continuity of the movement trajectory at the splicing of the front and rear materials in subsequent editing. For example: switch the camera at the mutation point and use the action vertex or micro-momentum connection to avoid a sense of jump.
[0098] Step S3. Construct a material relationship graph based on the similarity of the dynamic effect features of each video material.
[0099] That is to say, calculate the similarity between the dynamic effect feature data of each video material, and construct a material relationship graph with video materials as nodes and the similarity associations between video materials as edges.
[0100] The calculation of the similarity of the dynamic effect features of each video material includes the following steps:
[0101] Step S31. Convert the dynamic effect feature data of the video material into a numerical vector for calculating similarity, and calculate the similarity respectively:
[0102] Feature Vector = {trajectory similarity, speed similarity, direction similarity, position similarity}, specifically:
[0103] Trajectory similarity: Use dynamic time warping (DTW) to calculate the matching distance between two trajectories: Among them, T A , T B are the coordinate sequences of two trajectories, N is the trajectory length, and i represents the i-th point.
[0104] Speed similarity: Calculate the Euclidean distance of the speed sequence:
[0105] Among them, v A (t), v B (t) are the speed values of the two materials at the t-th frame.
[0106] Direction similarity: Calculate the cosine similarity of the direction vectors: DirSim(θ A , θ B ) = cos(θ A - θ B ), where θ A , θ B are the average direction angles of the two video materials;
[0107] Position similarity: Compare the coordinate distance between the end frame and the start frame:
[0108]
[0109] Furthermore, fuse the multi-feature similarities through weighted summation: Similarity(A, B) = w 轨迹 ·
[0110] DTW(T A , T B ) + w 速度 ·SpeedSim(SA , S B ) + w 方向 ·DirSim(θ A , θ B ) + w 位置 ·
[0111] DirSim(θ A , θ B ), where w 轨迹 , w 速度 , w 方向 , w 位置 are preset weights. In this embodiment, w 轨迹 = 0.4, and the rest are 0.2;
[0112] Normalize the similarity value to the [0, 1] interval to facilitate the subsequent construction and sorting of the material relationship graph.
[0113] Step S32: Set a node for each video material to store its dynamic effect feature vector: Node(i) = {ID, trajectory, speed sequence, direction sequence, position coordinates}; traverse all pairs of video materials, and use the similarity value calculated in step S31 as the edge: Edge(i, j) = {Node i , Node j , Similarity(i, j)}, and construct the initial relationship graph.
[0114] Step S33: Set a similarity threshold (such as 0.8, 0.6, 0.5, or 0.4), and remove the weak association edges with similarity lower than the threshold on the initial material relationship graph to obtain the material relationship graph.
[0115] The material relationship graph constructed by the present invention visualizes the similarity of the dynamic effect features (such as motion trajectory, speed, direction, position) of video materials as an association network through the graph structure of nodes and edges. Nodes represent individual video materials, edges represent the similarity strength between materials, and the weights of the edges quantify their association degree (such as the matching degree of trajectory, speed, and direction). It can replace manual frame-by-frame matching, especially suitable for the efficient processing of hundreds of materials. And based on the edge weights of the graph, the optimal timing line is generated through the path search algorithm to ensure the coherence of the spliced materials in terms of motion trajectory, speed, and direction. In short, the material relationship graph transforms complex dynamic effect features into a computable association network, providing the core decision-making basis for automated editing and significantly improving efficiency and creative quality.
[0116] Step S4: Generate a timing line according to the association degree between each video material on the material relationship graph.
[0117] That is to say, according to the edge weights between nodes in the material relationship graph, an optimal editing path is quickly generated among a large number of materials to ensure the maximum matching degree of adjacent video materials in terms of motion trajectory, speed, and direction.
[0118] As Figure 3 shown, the generation of the time sequence line includes the following steps:
[0119] Step S41: Randomly select several high centrality nodes as candidate starting points to generate an initial path pool. The high centrality nodes are the nodes with the highest PageRank value in the material relationship graph.
[0120] Step S42: Use the weighted longest path algorithm. For the end node of each path, traverse its unvisited neighbor nodes, expand the path in descending order of edge weights, and calculate the cumulative quality function Q of the path after each expansion. Only retain the M paths with the highest Q values;
[0121] Cumulative quality function Q: Among them, the weight coefficients (α = 0.5, β = 0.3, γ = 0.2), giving priority to ensuring the continuity of the motion trajectory.
[0122] Step S43: Stop expanding when the path covers all nodes or the correlation degree of the remaining nodes is lower than the threshold (such as <0.4), and generate several time sequence lines including the order of video materials.
[0123] Step S5: Sequentially splice the corresponding video materials according to the time sequence line, and customize the transition method between adjacent video materials to obtain the edited video.
[0124] That is to say, the time sequence line contains the order of video materials. The corresponding video materials are sequentially spliced according to the order, and the optimal transition method is dynamically selected for the transition between every two adjacent video materials, finally generating a video with natural and coherent visual effects.
[0125] It should be noted that according to the dynamic effect feature differences of adjacent video materials, the transition type is dynamically selected, specifically:
[0126] If the motion trajectory directions of adjacent video materials are continuous, use optical flow interpolation to generate a motion blur transition to simulate the effect of the camera following the moving object;
[0127] If the main body positions of adjacent video materials are complementary, generate a dynamic mask according to the positions and use coordinate interpolation to achieve spatial connection;
[0128] If the motion directions of adjacent video materials are opposite, insert a direction reversal special effect and use the principle of conservation of momentum to enhance the visual impact;
[0129] There are significant differences in the speeds of adjacent video materials. Insert a speed fade-in animation to achieve a smooth transition through a speed curve.
[0130] In step S5, the determination of the start frame and end frame of each video material includes the following steps:
[0131] Step S51: For each pair of adjacent video materials {A i , A i+1}, load the dynamic effect feature data of both.
[0132] Step S52: Through the key frame matching algorithm, screen out the candidate key frames strongly related to the motion trajectory from each video material, and determine the end frame of A i with the most matching motion features and the start frame of A i+1 between the candidate end frame of A i and the candidate start frame of A i+1 ;
[0133] The extraction of candidate key frames includes the following steps:
[0134] Step S521: Select the first frame tstart and the last frame tend of the motion trajectory;
[0135] Step S522: Calculate the speed difference Δv(t) using a sliding window, and mark the frames where Δv(t) > φ v ;
[0136] Step S523: Calculate the change in the direction angle between adjacent frames, and mark the frames where Δθ(t) > φ θ ;
[0137] Step S524: Summarize steps S521, S522, and S523 to obtain the candidate key frame set K = {(t, W total )}.
[0138] It should be noted that among the candidate end frame of A i and the candidate start frame of A i+1 , perform position matching, and select the candidate end frame and candidate start frame with the closest positions as the end frame of A i and the start frame of A i+1 .
[0139] Step S53: According to the start frame and end frame of each video material determined in step S52, clip the corresponding video material.
[0140] Through the above steps, aiming to make the main motion trajectories connect as much as possible, select the end frame of the previous video material and the start frame of the next video material to ensure the coherence of the video after the clipping process.
[0141] The present invention proposes an innovative method for processing film and television video clips. Through the automatic extraction and analysis of dynamic effect features (including motion trajectories, speeds, directions, and positions), and in combination with a material relationship graph and a weighted path algorithm, it realizes the full-process automatic operation from material selection to the output of the final edited product. This method can not only quickly generate the optimal editing order, greatly shortening the time required to process a large number of materials and significantly improving the overall efficiency of video editing, but also maintain a high degree of professionalism and accuracy during the processing.
[0142] In addition, the present invention also intelligently customizes transition effects through the dynamic matching of dynamic effect features between adjacent materials. For example, it adopts optical flow interpolation technology based on similar motion trajectories, or inserts fade-in and fade-out animations according to speed differences, etc., to ensure the natural and smooth transition between each frame of the picture. This method not only enhances the overall visual coherence of the video, but also improves its artistic expressiveness, providing a more immersive viewing experience for the audience. Therefore, the present invention is not only applicable to the professional film and television production field, but also provides strong technical support for self-media creators, advertisers, and other industries that require efficient and high-quality video production.
[0143] Exemplary electronic device
[0144] Figure 4 The block diagram of the electronic device according to an embodiment of the present application is shown.
[0145] As Figure 4 shown, the electronic device 10 includes one or more processors 11 and a memory 12.
[0146] The processor 11 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 10 to perform desired functions.
[0147] The memory 12 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 11 can run the program instructions to implement the film and television video clip processing method described above and / or other desired functions.
[0148] In one example, the electronic device 10 can further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0149] The input device 13 includes, for example, a keyboard, a mouse, and the like.
[0150] The output device 14 can output various information to the outside, including the video after clip processing and the like. The output device 14 can include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto, and the like.
[0151] Of course, for the sake of simplicity, Figure 4 only some of the components related to the present application in the electronic device 10 are shown, and components such as a bus, an input / output interface, and the like are omitted. In addition, according to specific application situations, the electronic device 10 may further include any other appropriate components.
[0152] Exemplary computer program product and computer storage medium
[0153] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the video and film clip processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0154] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0155] Furthermore, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the video and film clip processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0156] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0157] Based on the inspiration of the ideal embodiments of the present invention as above, through the above description, relevant personnel can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and the technical scope must be determined according to the scope of the claims.
Claims
1. A film and television video editing processing method, characterized in that: The following steps are involved: Get multiple video clips; Extracting the motion effect features of each of the video materials; Based on the similarity of the motion effect features of each of the video materials, construct a material relationship map; Generate a timeline according to the degree of association between each of the video materials on the material relationship graph; The corresponding video materials are spliced in sequence according to the timing line, and a transition method is customized between adjacent video materials to obtain a clipped video.
2. The method for film and television video editing according to claim 1, characterized in that: The method for extracting the dynamic effect feature comprises the following steps: Preprocessing the video material, including grayscale conversion, noise reduction and smoothing, and size adjustment; Calculate the optical flow field of the video material using the Lucas-Kanade algorithm; A motion effect feature is extracted from the optical flow field.
3. The method for video editing according to claim 2, characterized in that: The motion effect features include: the motion trajectory, speed, direction and position of the subject in the video screen; Motion trajectory: Track feature points based on the Lucas-Kanade algorithm and save the coordinate sequence of each point; speed: Among them, u and v are the components of the optical flow vector; direction: Position: Select the coordinate sequence of several points in the motion trajectory, including at least the coordinate sequences at both ends of the motion trajectory.
4. The method for video editing according to claim 3, characterized in that: The similarity calculation between the motion effect features of the video materials includes the following steps: The dynamic feature data of the video material is converted into a numerical vector for calculating similarity, and the similarity is calculated respectively: Feature Vector = {trajectory similarity, speed similarity, direction similarity, position similarity}, specifically: Trajectory similarity: Dynamic Time Warping (DTW) is used to calculate the matching distance between two trajectories: Among them, T A , T B is the coordinate sequence of the two trajectories, N is the trajectory length, and i is the i-th point on the motion trajectory; Speed similarity: Calculate the Euclidean distance of the speed sequence: Among them, v A (t), v B (t) is the velocity value of the tth frame of the two materials; Directional similarity: Calculate the cosine similarity of directional vectors: DirSim(θ A ,θ B )=cos(θ A -θ B ), where θ A ,θ B is the average direction angle of the two video materials; Position similarity: Compare the coordinate distance between the end frame and the start frame: Furthermore, the similarity of multiple features is fused by weighted summation: Similarity(A,B)=w 轨迹 · DTW(T A ,T B )+w 速度 ·SpeedSim(S A ,S B )+w 方向 ·DirSim(θ A ,θ B )+w 位置 · DirSim(θ A ,θ B ), where w 轨迹 、w 速度 、w 方向 、w 位置 is the preset weight; Normalize the similarity value to the interval [0,1].
5. The method for film and television video editing according to claim 2, characterized in that: The method for constructing the material relationship map includes: A node is set corresponding to each video material to store the dynamic effect feature vector: Node(i) = {ID, trajectory, speed sequence, direction sequence, position coordinates}; Traverse all the video material pairs, using the similarity value as the edge: Edge(i,j) = {Node i ,Node j , Similarity(i,j)}, construct the initial relationship graph; A similarity threshold is set, and weakly associated edges with similarities lower than the threshold on the initial material relationship graph are removed to obtain a material relationship graph.
6. The method for video editing according to claim 1, characterized in that: The generation of the timing line comprises the following steps: Randomly select several high centrality nodes as candidate starting points to generate an initial path pool; The weighted longest path algorithm is used. For each end node of each path, its unvisited neighbor nodes are traversed, and the path is expanded in descending order of edge weights. After each expansion, the cumulative quality function Q of the path is calculated, and only the M paths with the highest Q values are retained. Cumulative mass function Q: Among them, the weight coefficients α = 0.5, β = 0.3, γ = 0.2; When the path covers all nodes or the correlation degree of the remaining nodes is lower than a threshold, the expansion is stopped, and a plurality of time series lines containing the sequence of the video materials are generated.
7. The method for video editing according to claim 3, characterized in that: According to the difference in the dynamic effect characteristics of the adjacent video materials, the transition type is dynamically selected, specifically: The motion trajectories of adjacent video materials are continuous in direction, and optical flow interpolation is used to generate motion blur transition; The positions of the adjacent video material bodies are complementary, and a dynamic mask is generated according to the positions; The adjacent video materials move in opposite directions, and a direction reversal special effect is inserted; The speeds of adjacent video materials are significantly different, and a speed gradient animation is inserted.
8. The method for video editing according to claim 1, characterized in that: The determination of the start frame and the end frame of each video material includes the following steps: For each pair of adjacent video materials {A i , A i+1 }, load the dynamic feature data of both; Through the key frame matching algorithm, candidate key frames that are strongly related to the motion trajectory are selected from each of the video materials and i The candidate end frame and A i+1 The candidate start frames are determined to determine the A with the best motion feature match. i The end frame and A i+1 The start frame of According to the determined start frame and end frame of each of the video materials, the corresponding video materials are clipped.
9. An electronic device, characterized in that: include: processor; Memory; as well as The computer program instructions stored in the memory, when executed by the processor, enable the processor to execute the film and television video editing processing method according to any one of claims 1-8.
10. A computer storage medium, characterized in that It includes computer program instructions, which, when executed by a processor, enable the processor to execute the film and television video editing processing method as described in any one of claims 1-8.
Citation Information
Cited By
Video processing method and device, equipment and storage medium
CN121392692A