Short video editing method

By constructing a three-dimensional scene spatial model and calculating the viewpoint difference matrix, quantifying the viewpoint difference and optimizing the viewpoint switching, the problem of lack of logic in viewpoint selection and transformation in multi-view video clips is solved, and the quality and coherence of short videos are improved.

CN120390120APending Publication Date: 2025-07-29ZHENGZHOU INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510513226.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art lacks objective quantitative analysis of camera spatial relationships in multi-view video editing, resulting in a lack of logic and coherence in viewing angle selection and conversion, affecting the viewing experience of short videos.

Method used

By extracting the depth of field feature data and calculating the view angle difference matrix, a three-dimensional scene spatial model is constructed, and view angle differences are quantified using deep learning and feature matching algorithms, and view angle complementarity calculation is introduced to optimize view angle switching.

Benefits of technology

It realizes organic collaborative editing of multi-view videos, ensuring the logic and coherence of viewing angle switching, and improving the quality and viewing experience of short videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390120A_ABST
    Figure CN120390120A_ABST
Patent Text Reader

Abstract

The invention discloses a short video editing method, and relates to video editing, and the method comprises the steps: obtaining a plurality of video streams associated with a target scene, and the position data and shooting parameters of each camera device; constructing a three-dimensional scene space model according to the position data and the shooting parameters of each camera device; extracting depth-of-field feature data and scene position data of the multiple segments of video streams; obtaining a video frame data set; calculating a visual angle difference matrix of each camera device according to the depth-of-field feature data and the scene position data; performing segmentation processing on the first video stream by using the video frame data set to obtain a first video sequence; and editing the second video stream by using the video frame data set and the view angle difference matrix according to the first video sequence to obtain a second video sequence as an edited short video. As the multi-view video content lacks spatial perception integration, the short video quality is improved by extracting the depth-of-field feature data and calculating the view difference matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video editing, and in particular to a short video editing method. Background Art

[0002] With the rapid development of mobile internet and social media, short videos have become one of the most mainstream content formats today. Scenes such as sporting events, concerts, and film shoots often utilize multiple cameras simultaneously to capture content from different angles. These multi-view video resources provide a rich source of material for short video creation, allowing works to present the same scene from multiple perspectives, enhancing the audience's immersion and visual experience. Creating multi-view short videos requires selecting high-quality clips from a large amount of raw video material and then carefully editing and combining them. High-quality multi-view short videos should feature smooth perspective transitions, coherent narrative content, and visual harmony, which places high demands on editing techniques.

[0003] Traditional multi-view short video editing relies heavily on the experience and intuition of video editors to select footage and design perspective switching. Professionals typically need to review all footage, manually label high-quality segments, and then stitch them together based on content relevance and visual aesthetics. This approach has significant drawbacks: the same scene captured by cameras in different locations can exhibit significant variations in composition, subject prominence, and spatial relationships due to differing perspectives. Existing technologies lack objective quantitative analysis of these spatial relationships, making it difficult to achieve optimal perspective selection and switching.

[0004] For example, the related patent document CN119676510A discloses a method, device, and electronic device for automated video editing. The method includes: obtaining multiple video streams associated with a target scene; performing timestamp matching on the video frames included in each of the multiple video streams based on the key frame sequences corresponding to each of the multiple video streams to obtain a video frame mapping information set, wherein the video frame mapping information set includes multiple natural timestamps and multiple video frames that match the multiple natural timestamps, each corresponding to a video stream; upon obtaining a first video sequence obtained by clipping a first video stream from the multiple video streams, synchronously clipping at least one second video stream from the multiple video streams based on the video frame mapping information set to obtain a second video sequence that matches the first video sequence. However, this method relies solely on timestamp matching for editing and does not consider the spatial positional relationship and shooting parameters of different camera devices, resulting in an insufficient understanding of the spatial structure of the scene. Summary of the Invention

[0005] In view of the lack of spatial perception integration in multi-view video content in the prior art, the present application provides a short video editing method, which improves the quality of short videos by extracting depth-of-field feature data and calculating the perspective difference matrix.

[0006] The present application provides a short video editing method, including: S1, obtaining multiple video streams associated with the target scene, as well as the position data and shooting parameters of each camera device; the multiple video streams include a first video stream and a second video stream, and the first video stream represents the main perspective video data used as the editing reference; the second video stream represents the auxiliary perspective video data; S2, constructing a three-dimensional scene space model according to the position data and shooting parameters of each camera device; S3, extracting the depth-of-field feature data and scene position data of the multiple video streams; and performing time synchronization on the video frames of the multiple video streams to obtain a video frame data set; S4, calculating the perspective difference matrix of each camera device according to the depth-of-field feature data and scene position data; S5, performing segmentation processing on the first video stream by using the video frame data set to obtain a first video sequence; S6, editing the second video stream by using the video frame data set and the perspective difference matrix according to the first video sequence to obtain a second video sequence as the edited short video.

[0007] Further, S2, constructing a three-dimensional scene space model includes: extracting spatial coordinate data from the position data of each camera device, and extracting lens focal length, field of view angle, and direction vector data according to the shooting parameters; establishing a spatial position distribution model of the camera devices according to the spatial coordinate data, and calculating the shooting range projection data of each camera device according to the lens focal length, field of view angle, and direction vector data; constructing a three-dimensional scene space model of the target scene according to the spatial position distribution model and the shooting range projection data.

[0008] Further, S3, extracting the depth-of-field feature data and scene position data of the multiple video streams includes: performing frame decomposition processing on the multiple video streams to generate an initial video frame set; using a deep learning model to perform image analysis on the initial video frame set to identify the clear area and blurred area of each video frame, and generating a clarity heat map data; the clear area represents the focus area with clear imaging in the image, and the blurred area represents the non-focus area with blurred imaging in the image; calculating the depth-of-field feature data of each video frame according to the clarity heat map data, and the depth-of-field feature data represents the focal depth distribution of different areas in the video frame; performing object detection on the initial video frame set to identify the scene main body in the video frame, and the scene main body represents the main shooting object in the video; mapping the scene main body to the three-dimensional scene space model and calculating the position data of the scene main body in the three-dimensional scene space model as the scene position data.

[0009] Furthermore, the video frames of the multiple video streams are time-synchronized to obtain a video frame dataset, including: extracting the timestamps of the multiple video streams, and establishing the initial time relationship of the multiple video streams according to the timestamps; matching the video frames of the same scene subject in different video streams according to the initial time relationship and combining the scene position data; and performing time interpolation on the multiple video streams according to the matching results to obtain a time-calibrated video frame dataset.

[0010] Furthermore, S4 calculates the perspective difference matrix of each camera device, including: comparing and analyzing the clarity heat map data of each video frame in the first video stream and the second video stream based on the depth of field feature data, and generating video frame clarity difference data; mapping the scene position data into a three-dimensional scene space model, and calculating the spatial shooting angle difference data of different cameras for the same scene subject; calculating the content similarity data between the first video stream and the second video stream through a weighted fusion algorithm based on the video frame clarity difference data and the spatial shooting angle difference data; and constructing a perspective difference matrix reflecting the degree of perspective difference between each camera device based on the content similarity data.

[0011] Furthermore, generating video frame definition difference data includes: using depth of field feature data to extract definition heat map data of k pairs of video frames at corresponding times of the first video stream and the second video stream, and obtaining a paired heat map data set H1 = {H 11 ,H 12 ,....,H 1k} and H2={H 21 ,H 22 ,....,H 2k}; For each pair of video frames, the heat map data (H 1j ,H 2j ), and use bilinear interpolation or bicubic interpolation to obtain a data set H with the same resolution 1j ' and H 2j '; According to the three-dimensional scene space model, using the internal and external parameters of the camera, construct the space mapping function M j ; Using function M j H 2j 'Map to H 1j 'The corresponding spatial position, get H 2j ”; Calculate H 1j 'With H 2j "The pixel-level difference between |H 1j '(x,y)-H 2j ”(x,y)|, get the difference matrix D 1j , where H 1j '(x,y) represents the clarity value of the first video stream normalized heat map at coordinate (x,y), H 2j"(x,y) shows the clarity value of the mapped second video stream heat map at the same coordinates; calculate each difference matrix D 1j The average value μ j , get the single frame definition difference value; set the definition difference value μ of all video frame pairs j Composition difference dataset μ={μ1,μ2,.....,μ k}, as the video frame clarity difference data.

[0012] Further, calculating the spatial shooting angle difference data of the same scene subject by different camera devices includes: determining the coordinate point set of the scene subject in the three-dimensional scene space according to the scene position data obtained in step S3

[0013] P={p1,p2,.....,p n}; Using the three-dimensional scene space model, extract the spatial position coordinate set C1 of m pairs of camera devices corresponding to each frame in the first video stream and the second video stream = {C 11 ,C 12 ,.....,C 1m} and C2={C 21 ,C 22 ,.....,C 2m}; For each pair of camera positions (C 1i ,C 2i ), calculate from position C 1i and C 2i To the center point p of the scene's main coordinate point set P c The direction vector V 1i and V 2i ; From the video frame data set obtained in step S3, for each pair of camera positions (C 1i ,C 2i ) corresponding to the video frame, extract the image area containing the scene subject; for each pair of camera device positions (C 1i ,C 2i ) corresponds to the main image area of the scene, and the feature point set F is extracted using SIFT or ORB algorithm 1i and F 2i ; For each pair of feature point sets (F 1i ,F 2i ), calculate the feature point matching rate Among them, N i,match is the number of successfully matched feature point pairs; at the same time, the deformation matrix T describing the spatial mapping relationship of the feature points is calculated i , Among them, x1 and x2 are F 1i and F 2i The homogeneous coordinates of the corresponding feature points in ; calculate each pair of direction vectors V 1iand V 2i The angle θ between i For each pair of camera positions (C 1i ,C 2i ), according to the feature point matching rate R i , deformation matrix T i The eigenvalue and angle θ i , calculate the spatial shooting angle difference value D i : Among them, w1, w2, w3 are weight coefficients and satisfy σ i,max and σ i,min They are the deformation matrices T i The maximum and minimum non-zero singular values of all the camera device pairs’ spatial shooting angle difference values D i Composition difference data set D = {D1, D2, ..., D m}, as the spatial shooting angle difference data of different camera devices on the same scene subject.

[0014] Furthermore, the content similarity data between the first video stream and the second video stream is calculated by a weighted fusion algorithm, including: establishing a video frame definition difference data set μ={μ1,μ2,......,μ k} and the spatial shooting angle difference dataset D={D1,D2,......,D m}, and obtain the paired difference dataset (μ i ',D i '), i = 1, 2, ..., q, where q is the number of pairs; for each pair of difference data (μ i ',D i '), calculate the content similarity value S i :S i =1-[α×norm(μ i ')+β×norm(D i ')], where α and β are weight coefficients and satisfy α+β=1, norm(*) is a normalization function that maps the difference value to the interval [0, 1]; the content similarity value sequence {S1, S2, ..., S q} as content similarity data between the first video stream and the second video stream.

[0015] Furthermore, constructing a perspective difference matrix reflecting the degree of perspective difference between the camera devices includes: converting the content similarity value sequence {S1, S2, ..., S qDivide it into n subsequences according to time periods, and each subsequence corresponds to a scene time interval; calculate the statistical feature values of each subsequence as the similarity feature vector SV corresponding to the time interval i ; According to the similarity feature vector SV i , calculate the viewing angle difference degree d between the first video stream and the second video stream in each time interval i , d i = g(SV i ), where the mapping function

[0016] d i = g(SV i ) = 1 - (w1×mean(SV i ) + w2×(1 - var(SV i )) + w3×(1 - max(SV i )) + w4×min(SV i ))), w1, w2, w3, w4 are weight coefficients and satisfy w1 + w2 + w3 + w4 = 1; mean(SV i ) represents the mean value of SV i ; var(SV i ) represents the variance of SV i ; max(SV i ) represents the maximum value of SV i ; min(SV i ) represents the minimum value of SV i ; Construct an n×n - dimensional viewing angle difference matrix M, where the matrix element M ij represents the viewing angle jump degree between the video content in the i - th time interval and the video content in the j - th time interval, and the calculation formula is: where, t i and t j are the central moments of the i - th and j - th time intervals respectively.

[0017] Further, in S5, obtain the first video sequence, including: perform scene content analysis on the first video stream, extract scene change features according to depth - of - field feature data and scene position data; according to the scene change features, use a threshold detection algorithm to identify scene switching points in the first video stream, and divide the first video stream into multiple initial scene segments {SP1, SP2,......, SP l}; calculate the content score value E i of each initial scene segment SP i , E i = λ1×C i +λ2×S i , C i represents the clarity score, HLj Indicates the segment SP i The clarity heat map of the j-th frame in, n i Is the segment SP i The number of frames of, max(HL j ) represents the maximum clarity value of the scene main area; S i Represents the stability score, Indicates the segment SP i The variance of the clarity difference value between consecutive frames in; Select the segment with the content score value E i Greater than the threshold as the candidate segment, to obtain the first video sequence SEQ1, SEQ1 = {SEQ 11 , SEQ 12 ,......, SEQ 1r}, and its corresponding time index information T1, T1 = {T 11 , T 12 ,....., T 1r}.

[0018] Furthermore, in S6, obtain the second video sequence as the short video for editing, including: According to the time index information T1 of the first video sequence SEQ1, find the corresponding second video stream paragraph set US2 = {US 21 , US 22 ,....., US 2r} in the video frame dataset; Use the perspective difference matrix M to calculate the perspective complementarity VC between each scene segment in the first video sequence SEQ1 and the corresponding time period of the second video stream, VC = γ × M ij +(1 - γ)×(1 - S i ), where, M ij Is the corresponding element of the perspective difference matrix, S i Is the content similarity value of the corresponding time period, and γ is the weight coefficient; Select segments with a perspective difference greater than the threshold from the second video stream according to the perspective complementarity VC to form the primary selection segment set US2'; Combine the primary selection segment set US2' in chronological order to obtain the second video sequence SEQ2 as the short video for editing.

[0019] Compared with the prior art, the advantages of this application are as follows:

[0020] (1) On the one hand, in a multi-camera scenario, the content of different perspectives needs to be organically coordinated, but the prior art is difficult to understand the spatial relationship between cameras, resulting in a lack of logic and coherence in perspective switching. Traditional methods such as simple time slicing cannot perceive the spatial correlation of content from different perspectives, while manual editing is time-consuming, laborious and subjective.

[0021] This application extracts spatial coordinates from the position data of the imaging device, and extracts the lens focal length, field of view angle, and direction vector data from the shooting parameters; establishes a spatial position distribution model and shooting range projection data to form a complete three-dimensional scene understanding, which enables the system to place all cameras in the same coordinate system and understand the relative position relationship between them. Calculate the clarity difference of video frames by comparing the clarity heat map; use the SIFT / ORB algorithm to extract the feature point sets F 1i and F 2i , calculate the feature point matching rate and the deformation matrix; by calculating accurately quantify the viewing angle difference; construct an n×n-dimensional viewing angle difference matrix to represent the degree of viewing angle jump between the contents in different time intervals.

[0022] (2) On the other hand, the lack of coherence in viewing angle switching is a common problem in short video editing, which will lead to a fragmented viewing experience. Existing technologies mostly use fixed-duration switching or matching based on simple features, and cannot ensure the natural and smooth visual transition. This application introduces the viewing angle complementarity VC = γ×M ij +(1 - γ)×(1 - S i ); combine the viewing angle difference matrix M and the content similarity S i for weighted calculation; balance the dual requirements of viewing angle difference and content relevance. Find the corresponding segment in the second video stream according to the time index of the first video sequence; select the segment with a viewing angle complementarity greater than the threshold to ensure that the selected viewing angles are both different and related; combine them in chronological order to form the final editing result. Through the scientific calculation of the viewing angle complementarity, the system can find the best viewing angle switching point. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] This application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive, and in these embodiments, the same numbers represent the same structures, where:

[0024] Figure 1 is an exemplary flowchart of a short video editing method shown in some embodiments of this application;

[0025] Figure 2 is a schematic diagram of the intelligent editing process of multiple camera viewing angles shown in some embodiments of this application;

[0026] Figure 3 is a schematic diagram of the position of the imaging device shown in some embodiments of this application;

[0027] Figure 4 is a schematic diagram of the co-viewing area shown in some embodiments of this application;

[0028] Figure 5is the clarity heat map shown in some embodiments of the present application;

[0029] Figure 6 is the video frame time synchronization effect diagram shown in some embodiments of the present application;

[0030] Figure 7 is the heat map H shown in some embodiments of the present application 1j (7a) and the standardized heat map H 1j '(b);

[0031] Figure 8 is the heat map H shown in some embodiments of the present application 2j (8a) and the standardized heat map H 2j '(8b);

[0032] Figure 9 is the score distribution schematic diagram shown in some embodiments of the present application;

[0033] Figure 10 is the schematic diagram of the paragraph set US2 shown in some embodiments of the present application. Detailed implementation manners

[0034] The methods and systems provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0035] As Figure 1 and Figure 2 shown, S1, obtain multiple video streams associated with the target scene, as well as the position data and shooting parameters of each camera device; the multiple video streams include a first video stream and a second video stream, the first video stream represents the main perspective video data as the editing reference; the second video stream represents the auxiliary perspective video data; S2, construct a three-dimensional scene space model according to the position data and shooting parameters of each camera device; S3, extract the depth-of-field feature data and scene position data of the multiple video streams; and perform time synchronization on the video frames of the multiple video streams to obtain a video frame data set; S4, calculate the viewing angle difference matrix of each camera device according to the depth-of-field feature data and scene position data; S5, perform segmentation processing on the first video stream by using the video frame data set to obtain a first video sequence; S6, according to the first video sequence, use the video frame data set and the viewing angle difference matrix to edit the second video stream to obtain a second video sequence as the edited short video.

[0036] S1, in this embodiment, first, it is necessary to obtain multiple video streams associated with the target scene and related data. The system obtains video data captured by multiple camera devices synchronously or asynchronously in the target scene through a video acquisition interface. These videos can come from multiple fixed cameras, mobile cameras or handheld devices and are stored in the original video file format (such as MP4, MOV, etc.).

[0037] The system divides the acquired multiple video streams into a first video stream and a second video stream. Among them, the first video stream is designated as the main perspective video data for the editing reference, usually selecting the video stream with the most comprehensive shooting perspective, the best picture quality, or the most prominent subject performance; the second video stream serves as the auxiliary perspective video data, providing supplementary pictures from multiple angles.

[0038] The system obtains the position data of each camera device in the following ways: for camera devices with GPS positioning functions, directly read their geographical location coordinates; for fixed camera devices, obtain through pre-measured and recorded position data; for mobile camera devices, their relative positions can be deduced through visual SLAM (Simultaneous Localization and Mapping) technology or image-based position estimation algorithms.

[0039] The system obtains the shooting parameters of each camera device by reading the metadata of the video file or device recording information, including: lens focal length: a parameter representing the optical characteristics of the lens, usually in millimeters; aperture size: controlling the amount of light entering and the depth of field effect; field of view: the angle representing the shooting range of the camera device; direction information: the orientation of the camera device, which can be estimated through gyroscope data or from image features.

[0040] S2. According to the position data and shooting parameters of each camera device, construct a three-dimensional scene space model. Specifically, extract the three-dimensional space coordinates of X, Y, and Z from the position data of the camera device, and establish a unified coordinate system; analyze the shooting parameter data and extract the lens focal length value of each camera device (usually ranging from 8mm to 200mm); calculate the field of view according to the focal length, and the typical field of view ranges from wide-angle (about 100°) to telephoto (about 10°); calculate the direction vector of each camera device, which is used to represent the orientation of the camera. In this embodiment, the positions of the camera devices are as Figure 3 shown. Among them, Camera #1 (main position), position: (0.0, 1.8, 0.0) m; focal length: 35mm | field of view: 63°; direction: (0.0, 0.0, 1.0). Camera #2 (side position), position: (3.5, 2.1, 4.2) m; focal length: 24mm | field of view: 84°; direction: (-0.4, -0.1, 0.5). Camera #3 (distant view), position: (-2.8, 2.5, 8.0) m; focal length: 85mm | field of view: 28°; direction: (0.2, -0.3, -0.8). Camera #4 (close-up), position: (1.2, 1.5, 2.5) m; focal length: 135mm | field of view: 18°; direction: (-0.1, 0.0, 0.2).

[0041] Using the extracted spatial coordinates, a spatial position distribution model is constructed to describe the relative position relationship of each camera device in the three-dimensional space. Based on the field of view angle and direction vector, the view frustum of each camera device is calculated to represent its shooting range. The view frustum is projected into the three-dimensional space to generate shooting range projection data, which is usually expressed as a spatial area in the shape of a quadrangular pyramid. The overlapping area of the shooting ranges of different cameras is calculated to determine the common viewing area of the scene. In this embodiment, the common viewing area is as follows: Figure 4 shown.

[0042] The spatial position distribution model and shooting range projection data are integrated to construct a complete 3D scene spatial geometry model. SfM (Structure from Motion) technology is used to extract feature points from multi-view images and calculate the 3D position of objects in the scene through triangulation. Multi-view geometry principles are applied to optimize 3D scene point cloud data and improve model accuracy. The position of each camera device and its viewing cone are marked in 3D space to form an intuitive 3D scene spatial model for subsequent analysis of the perspective relationship and content differences between the cameras.

[0043] S3 extracts depth feature data and scene position data from multiple video streams. In this embodiment, the video frame decomposition process uses an adaptive sampling strategy, specifically implemented as follows: Decoding method: The FFmpeg library (version 4.4+) is used as the underlying decoding engine, supporting mainstream codecs such as H.264 / H.265 / VP9 / AV1. For videos with resolutions of 1080p and below, full decoding is used; for 4K videos, they are downsampled to 1080p before processing to balance performance and accuracy. Sampling strategy: The system sets a base sampling rate of 24fps. For clips with drastic content changes (when the difference exceeds a threshold of 15% through continuous frame difference calculation), the sampling rate is increased to the original video frame rate (up to 60fps). For static scenes (frame difference less than 5%), the sampling rate is reduced to 8fps to reduce redundant calculations. Frame extraction parameters: Each frame is stored in the RGB color space with a 24-bit color depth and lossless format. For HDR videos, the system performs color gamut mapping to the standard sRGB space to ensure consistency in subsequent processing. Preprocessing optimization: Gaussian noise reduction (σ = 0.5-1.0, depending on the noise level) and grayscale histogram equalization are applied to the extracted frames to enhance image details and improve the accuracy of subsequent analysis.

[0044] like Figure 5As shown in the figure, the system of this embodiment uses a lightweight model that combines the improved DeblurGAN-v2 architecture with MobileNetV3 for clarity analysis: Network architecture: The backbone network uses MobileNetV3-Large, with an input size of 224×224×3. The feature extraction part contains 16 depthwise separable convolutional layers, and the number of parameters is about 5.4M; an FPN (Feature Pyramid Network) structure is added to achieve multi-scale feature fusion, with a total of 4 scale levels. Training data: The model is trained on a dataset containing 25,000 pairs of clear-blurry image pairs, covering image samples under different depths of field and different lighting conditions. The training uses the Adam optimizer, with an initial learning rate of 3×10^-4, a batch size of 32, and 100 training rounds. Clarity evaluation: The model outputs a single-channel heat map with a resolution of 256×256, and the pixel value range is [0, 1], where 1 represents the highest clarity and 0 represents the lowest clarity. The system uses a sliding window (32×32, step size 16) to perform inference on the original resolution, generating a heat map with the same resolution as the original image. Performance indicators: The model reaches an average MSE (Mean Squared Error) of 0.042 and an SSIM (Structural Similarity) of 0.89 on the validation set; the single-frame processing time is 15ms on an RTX 3080 and does not exceed 120ms on a mobile SoC.

[0045] Based on the clarity heat map, the system calculates the depth-of-field feature data: Focal plane positioning: The system locates the area with the highest clarity value (usually the threshold is set above 0.85) as the focal plane position and records its average depth value d focus (relative value, range [0, 1]). Depth-of-field range calculation: By analyzing the gradient change of clarity from high to low, the area where the clarity exceeds 70% is determined as the "acceptable clear area", and its nearest point depth d near and the farthest point depth d far are recorded. Depth-of-field distribution representation: The system divides the image into a 10×10 grid, and each grid calculates the average clarity value and the estimated depth value to form a depth-of-field distribution matrix SD. For each clarity cluster (adjacent high-clarity areas), record its center position (x, y), coverage radius r, average clarity value c, and estimated depth value d, generating a depth-of-field feature vector SDF = [x, y, r, c, d]. Depth-of-field feature parameterization: The final depth-of-field feature data is represented as: {d focus , d near , d far , DoF, SD, [SDF1, SDF2,......]}, where DoF = d far -d near represents the depth of the depth of field. Typical value range: DoF ∈ [0.05, 0.4], depending on the lens aperture and focal length.

[0046] The system uses a combination of YOLOv5s and DeepSORT for target detection and tracking. The specific implementation is as follows: Detection model configuration: The YOLOv5s model (with approximately 7.2M parameters) is used, with an input resolution of 640×640, supporting detection of 80 general object categories. Specialized fine-tuning models are added for specific scenarios (such as sports and concerts). The detection threshold is set to confidence > 0.45, and the NMS threshold is 0.4. Scene subject determination rules: The system determines the scene subject based on a weighted score of multiple factors: Size Factor S: the proportion of the target bounding box area to the frame, with a weight of 0.3; Center Factor C: the distance from the target to the center of the frame, with a weight of 0.25; Clarity Factor F: the average clarity of the target area, with a weight of 0.3; and Duration Factor D: the proportion of the target's duration in the video, with a weight of 0.15. The scene subject score is calculated as follows: Score = 0.3 × S + 0.25 × C + 0.3 × F + 0.15 × D. The system selects objects with a score > 0.65 as scene subjects, retaining a maximum of three primary scene subjects. Tracking Implementation: Object tracking is performed using the DeepSORT algorithm, with a feature extraction network using the MobileNet backbone, an appearance feature dimension of 128, a matching threshold of 0.45, a maximum number of lost frames (max_age) of 30, and a minimum detection count (min_hits) of 3.

[0047] Based on the three-dimensional scene space model established in the previous steps, the system maps the scene subject to the three-dimensional space. The specific implementation is as follows: Single-view depth estimation: For a single video stream, the system applies the MiDaSv3.1 depth estimation network, inputs a resolution of 384×384, outputs a relative depth map, and calculates the initial depth estimate value of the scene subject.

[0048] Multi-view fusion mapping: Combining information from multiple video streams, the system performs the following steps: extracting key points of the scene subject from each perspective (such as 17 skeleton points of the human body or 8 corner points of an object); using the DLT (direct linear transform) algorithm to calculate the three-dimensional coordinates based on the camera parameters and the detected key points; applying the core algorithm RANSAC to filter outliers and improve triangulation accuracy; for occluded parts, using timing information and rigid body constraints for interpolation estimation; Position data representation: The three-dimensional position data of the scene subject is represented as: {object_id, type, [x, y, z], [w, h, d], orientation, confidence}; object_id: unique identifier of the scene subject; type: subject type (such as person, car, other object);

[0049] [x, y, z]: coordinates of the center of the three-dimensional space (in meters, relative to the scene origin); [w, h, d]: width, height, depth (in meters); orientation: quaternion [qx ,q y ,q z ,q w ; Confidence: Position estimation confidence (range [0, 1]); Precision control: The average 3D reconstruction precision reaches ±15 cm (within 5 meters), the confidence threshold is set to 0.7, and position data below this threshold is optimized and predicted through a Kalman filter.

[0050] Video frame time synchronization processing. The system extracts time information from multiple video streams and establishes an initial correspondence. The specific implementation is as follows: Timestamp data source: The system extracts time information from the following three levels: File metadata: Extract the absolute time starting point from the creation_time field of the MP4 / MOV container; Video stream metadata: Extract time code information from the SEI (Supplemental Enhancement Information) packet; Content time marker: Identify digital clocks or time displays that may appear in the video; Time unit and precision: The system uses millisecond-level time precision, internally represented as a 64-bit integer timestamp, with the starting point being the Unix epoch (January 1, 1970 UTC). For the effect of video frame time synchronization processing in this embodiment, see Figure 6 .

[0051] Based on the initial time relationship and scene position data, the system performs precise matching on frames in different video streams. The specific implementation is as follows: Feature point extraction and matching: The system uses the SIFT algorithm (extract 300 - 500 key points per frame) or the ORB algorithm (extract 800 - 1000 key points per frame) to extract feature points; For each pair of potentially corresponding frames, use FLANN (Fast Library for Approximate Nearest Neighbors) to perform feature matching, and the ratio test threshold is set to 0.75; Calculate the RANSAC homography matrix, and the inlier ratio threshold is set to 0.4.

[0052] Subject consistency verification: For the identified scene subjects, the system uses subject-specific matching strategies: Human subjects: Apply human pose estimation (OpenPose) to compare the spatial configuration similarity of 17 key points; Object subjects: Compare the aspect ratio of the bounding box, texture histogram (64 dimensions), and HOG features (324 dimensions); Similarity scoring function: Sim(A, B)

[0053] = 0.4×(Key point matching score)+0.3×(Pose / shape similarity score)+0.3×(Texture similarity score); The matching threshold is set to Sim(A, B)>0.65, and matches below this threshold are rejected.

[0054] Time Window Search: Based on the initial time relationship, the system searches for potentially matching frames within the range of ±1 second around the predicted time point; for each frame to be matched, the system evaluates up to 30 candidate frames at most. In descending order of similarity, the system adopts a scheme combining YOLOv5s and DeepSORT for object detection and tracking. The specific implementation is as follows: Detection Model Configuration: The YOLOv5s model (with approximately 7.2M parameters) is adopted, with an input resolution of 640×640, supporting the detection of 80 common objects; for specific scenarios (such as sports, concerts, etc.), a dedicated fine-tuned model is added. The detection threshold is set to confidence>0.45, and the NMS threshold is 0.4. Scene Subject Judgment Rule: The system determines the scene subject based on a multi-factor weighted score: Size Factor S: The proportion of the target bounding box area in the frame, with a weight of 0.3; Center Factor C: The distance from the target to the center of the frame, with a weight of 0.25; Clarity Factor F: The average clarity of the target area, with a weight of 0.3; Duration Factor D: The proportion of the target's duration in the video, with a weight of 0.15. Scene Subject Score Calculation Formula: Score = 0.3×S + 0.25×C + 0.3×F + 0.15×D. The system selects the target with Score>0.65 as the scene subject and retains at most 3 main scene subjects. Tracking Implementation: The DeepSORT algorithm is used for object tracking. The feature extraction network adopts a MobileNet backbone, with an appearance feature dimension of 128, a matching threshold of 0.45, a maximum number of lost frames (max_age) of 30, and a minimum detection count (min_hits) of 3.

[0055] Based on the three-dimensional scene space model established in the foregoing steps, the system maps the scene subject to the three-dimensional space. The specific implementation is as follows: Single-View Depth Estimation: For a single video stream, the system applies the MiDaSv3.1 depth estimation network, with an input resolution of 384×384, outputs a relative depth map, and calculates the initial depth estimation value of the scene subject.

[0056] Multi-View Fusion Mapping: Combining the information of multiple video streams, the system performs the following steps: Extract the key points of the scene subject from each perspective (such as 17 skeletal points of a human body or 8 corner points of an object); Based on the camera parameters and the detected key points, use the DLT (Direct Linear Transformation) algorithm to calculate the three-dimensional coordinates; Apply the core algorithm RANSAC to filter out outliers and improve the triangulation accuracy; For the occluded part, use temporal information and rigid body constraints for interpolation estimation; Position Data Representation: The three-dimensional position data of the scene subject is represented as: {object_id, type, [x, y, z], [w, h, d], orientation, confidence}; object_id: The unique identifier of the scene subject; type: The type of the subject (such as person, vehicle, other object);

[0057] [x, y, z]: Center coordinates in 3D space (in meters, relative to the scene origin); [w, h, d]: Width, height, depth (in meters); orientation: Orientation quaternion [q x , q y , q z , q w ; confidence: Confidence of position estimation (range [0, 1]); Precision control: The average 3D reconstruction precision reaches ±15 cm (within 5 meters), the confidence threshold is set to 0.7, and position data below this threshold is optimized and predicted through a Kalman filter.

[0058] Video stream synchronization point recognition: The system recognizes synchronization points with obvious features (such as sound peaks, flashes, action instants) as anchor points for initial time calibration. The synchronization point features detected by the system include: Instants when the sound intensity exceeds 3 times the average value; Frames with a change in picture brightness exceeding 40%; Key pose changes of the scene main body (such as clapping, jumping); Initial time relationship representation: Establish a time offset mapping table T offset , represented in the form of an NxN matrix (N is the number of cameras), and each element T offset [i][j] represents the time offset value (in milliseconds) from camera i to camera j, and the initial standard deviation of precision is set to ±200 ms.

[0059] Based on the initial time relationship and scene position data, the system performs precise matching on frames in different video streams, and the specific implementation is as follows: Feature point extraction and matching: The system uses the SIFT algorithm (extracting 300 - 500 key points per frame) or the ORB algorithm (extracting 800 - 1000 key points per frame) to extract feature points; For each pair of potentially corresponding frames, use FLANN (Fast Library for Approximate Nearest Neighbors) to perform feature matching, and the ratio test threshold is set to 0.75; Calculate the RANSAC homography matrix, and the inlier ratio threshold is set to 0.4.

[0060] Subject consistency verification: For the identified scene main body, the system uses a subject-specific matching strategy: Human subject: Apply human pose estimation (OpenPose) to compare the similarity of the spatial configuration of 17 key points; Object subject: Compare the aspect ratio of the bounding box, texture histogram (64 - dimensional), and HOG features (324 - dimensional); Similarity scoring function: Sim(A, B)

[0061] = 0.4×(Key point matching score)+0.3×(Pose / shape similarity score)+0.3×(Texture similarity score); The matching threshold is set to Sim(A, B)>0.65, and matches below this threshold are rejected.

[0062] Arrangement; for high-speed motion scenarios, the system expands the search window to ±2 seconds and increases the number of candidate frames to 50.

[0063] The matching result indicates that each successful frame match is recorded as:

[0064] {frame id,1 , frame id,2 , similarity score , confidence, transformation matrix}; frame id,x : The unique identifier of the frame in each video stream; similarity score : The similarity score, range [0, 1]; confidence: The confidence of the match, range [0, 1]; transformation matrix : A 3×3 transformation matrix that describes the spatial mapping relationship between two frames.

[0065] Based on the frame matching results, the system performs precise time synchronization calibration on multiple video streams as follows: Construction of the time mapping function: The system uses cubic spline interpolation to construct a continuous time mapping function; for each pair of video streams (i, j), a function T map,i,j (t) is established to map the time t of video stream i to the corresponding time of video stream j; the spline nodes are set to the time points of the matching frame pairs, usually one node is set every 10 to 15 seconds; the boundary conditions are set to natural boundary conditions (the second derivative is zero).

[0066] Anomaly detection and handling: The system calculates the residual standard deviation to identify abnormal matching points (deviation > 3σ); for the detected abnormal points, the system performs the following processing: If the abnormal point density < 5%, these points are directly removed; if the abnormal point density is 5% - 15%, the LOWESS (Locally Weighted Scatterplot Smoothing) algorithm is applied to reconstruct the time mapping function; if the abnormal point density > 15%, the video is segmented and a separate time mapping function is constructed for each segment.

[0067] Time calibration accuracy evaluation: The system uses 10% of the matching frame pairs as the validation set to calculate the time mapping error; an average time deviation < 16.7 ms (1 / 60 second) is considered high-precision synchronization; a deviation range of 16.7 ms - 33.3 ms is considered standard-precision synchronization; sections with a deviation > 33.3 ms are marked as low-precision synchronization and special processing strategies are applied in subsequent clips.

[0068] Video frame dataset generation: The system organizes the frames of multiple video streams into a unified video frame dataset according to the calibrated time mapping function; the dataset structure is designed as:

[0069] {time point , [frame1, frame2,....., frame n , main subject , depth features , sync quality}; time point : Standardized time point (milliseconds); frame x : Corresponding frames in each video stream, including ID and quality information; main subject : Scene main body information identified at this time point; depth features : Depth of field feature data for each perspective; sync quality : Synchronization quality score, range [0, 1].

[0070] Real-time synchronization support: The system supports an online learning mode and can dynamically update the time mapping function according to newly connected video stream data; the incremental update algorithm is executed every 5 seconds to ensure synchronization accuracy during long-term shooting; the time drift compensation mechanism applies Kalman filter prediction to handle possible device clock drift (usually <1ms / minute).

[0071] S4. According to the depth of field feature data and scene position data, combined with the visual parameter dataset, calculate the perspective difference matrix of each camera device. First, the system extracts the clarity heat maps of k pairs of video frames at the same time from the first video stream and the second video stream in the video frame dataset after time synchronization, forming paired datasets H1 = {H 11 , H 12 ,...., H 1k} and H2 = {H 21 , H 22 ,...., H 2k}. In practical applications, the value of k can be determined according to the video length, usually sampled at intervals of 0.5 - 1 second to ensure the continuity of analysis and calculation efficiency.

[0072] Since different video streams may have different resolutions, the system normalizes each pair of heat maps (H 1j , H 2j ). In actual implementation, for videos with a resolution of 720p and below, bilinear interpolation is used; for high-resolution videos of 1080p and above, a more computationally efficient bicubic interpolation method is used to process the two heat maps into the same standard resolution (usually 640×360 or 1280×720), obtaining the standardized heat map H 1j ' See Figure 7 , and the standardized heat map H 2j ', see Figure 8 .

[0073] Then, based on the three-dimensional scene space model constructed in step S2, the system constructs a spatial mapping function M using the internal parameters (such as focal length, principal point coordinates) and external parameters (position, orientation) of the camera. j . This function realizes the spatial transformation mapping from the perspective of the second video stream to the perspective of the first video stream.

[0074] The system transforms the standardized heat map H of the second video stream 2j ' through the mapping function M j to obtain H 2j ”, making it correspond to the heat map H of the first video stream in spatial position 1j ', that is, “aligning” the heat maps from two different perspectives to the same spatial reference system. In actual implementation, this process involves perspective transformation and resampling of images, and requires precise camera parameter calibration.

[0075] The system calculates the pixel-level difference between the aligned heat maps to generate a difference matrix D 1j , and its calculation formula is |H 1j '(x,y) - H 2j ”(x,y)|. This difference value quantifies the difference in clarity performance between the two perspectives at the same spatial position. Calculate the average value μ 1j for each difference matrix D j as the single-frame clarity difference value. In practical applications, the difference value can be weighted, and a higher weight can be assigned to the difference corresponding to the main area of the scene to improve the pertinence of the difference calculation. Combine the clarity difference values μ j of all k pairs of video frames to form a difference dataset μ = {μ1, μ2,....., μ k}, completing the generation of video frame clarity difference data. This dataset reflects the change in clarity performance difference between the two video streams over the entire time series.

[0076] Next, based on the scene position data obtained in step S3, the system determines the set of coordinate points P = {p1, p2,....., p n} of the scene main body in three-dimensional space. These coordinate points represent the contour or key point positions of the scene main body, and n is usually between 10 and 100, depending on the complexity of the scene main body.

[0077] The system extracts the spatial position coordinates of the corresponding camera devices of the first video stream and the second video stream at different times from the three-dimensional scene space model to form position sets C1 = {C 11 , C 12 ,....., C 1m} and C2 = {C 21 , C 22,.....,C 2m}. For fixed cameras, these positions remain unchanged; for moving cameras, these positions change over time.

[0078] The system calculates the direction vectors V 1i , C 2i ) pointing from each pair of camera positions (C c to the center point p of the scene subject 1i and V 2i . These vectors represent the "line-of-sight direction" of the cameras and are important parameters for evaluating the perspective difference. From the video frames, the system crops out the image region containing the scene subject, usually using the bounding box determined by the object detection result and appropriately expanding it to include sufficient context.

[0079] For each pair of corresponding scene subject image regions, the system applies the SIFT (for high precision requirements) or ORB (for high speed requirements) algorithm to extract the local feature point sets F 1i and F 2i . In practical applications, the number of feature points is usually controlled between 200 and 1000 to balance the computational efficiency and matching accuracy.

[0080] The system calculates the feature point matching rate where N i,match is the number of successfully matched feature point pairs. At the same time, the deformation matrix T i is calculated. This matrix describes the spatial mapping relationship between the two sets of feature points and reflects the degree of image deformation caused by the perspective difference. The system calculates the angle θ 1i between the direction vectors V 2i and V i , which directly reflects the perspective difference between the two cameras relative to the scene subject.

[0081] Finally, combining the feature point matching rate R i , the eigenvalue analysis of the deformation matrix, and the perspective angle θ i , the system calculates the comprehensive spatial shooting angle difference value D i : where: w1, w2, w3 are weight coefficients, adjusted according to the actual application. In this embodiment, the values are w1 = 0.3, w2 = 0.3, w3 = 0.4;

[0082] The degree of deformation is evaluated through the singular value ratio of the deformation matrix T i ; σ i,max , σ i,min are the maximum and minimum non-zero singular values of T i respectively.

[0083] For all pairs of cameras, the spatial shooting angle difference values Di Construct a differential dataset D = {D1, D2,......, D m}, as the spatial shooting angle difference data. This dataset quantifies the spatial difference degree of the camera view angle at different times.

[0084] The system aligns the video frame clarity difference dataset μ and the spatial shooting angle difference dataset D in time to form a paired difference dataset (μ i ', D i '), i = 1, 2,...., q. Since the sampling rates of the two types of difference data may be different, the system ensures data alignment through time interpolation method.

[0085] The system normalizes the paired difference data to ensure that different types of difference data are compared on the same magnitude: norm(μ i ') maps the clarity difference value to the interval [0, 1]; norm(D i ') maps the spatial angle difference value to the interval [0, 1]. For each pair of paired data, the system calculates the content similarity value S i :

[0086] S i = 1 - [α × norm(μ i ') + β × norm(D i ')], where: α and β are weight coefficients, satisfying α + β = 1, adjusted according to the scene characteristics. In a scene with rich content details, the value of α can be increased; in a scene with complex spatial structure, the value of β can be increased; the typical configuration is α = 0.4, β = 0.6. The calculated content similarity values are formed into a sequence {S1, S2,......, S q}, as the content similarity data between the first video stream and the second video stream. This sequence describes the change of the content similarity degree between the two video streams on the entire timeline.

[0087] Finally, construct a view angle difference matrix representing the view angle conversion difficulty between different time intervals. The system divides the content similarity value sequence {S1, S2,......, S q} into n subsequences in chronological order, and each subsequence corresponds to a scene time interval. The division can be based on the scene switching point, or use a uniform division with a fixed duration (such as 2 - 5 seconds).

[0088] Calculate the statistical feature values for each subsequence, including the mean mean(SV i ), variance var(SV i ), maximum value max(SV i ), minimum value min(SV i ), etc., to form the similarity feature vector SV corresponding to the time intervali According to the similarity feature vector SV i , calculate the perspective difference degree d of each time interval i :

[0089] d i = g(SV i ) = 1 - (w1×mean(SV i ) + w2×(1 - var(SV i )) + w3×(1 - max(SV i )) + w4×min(SV i ))), where: w1, w2, w3, w4 are weight coefficients, satisfying w1 + w2 + w3 + w4 = 1; the typical configuration is w1 = 0.5, w2 = 0.2, w3 = 0.2, w4 = 0.1; this formula comprehensively considers the average level, stability and extreme value situations of similarity.

[0090] The system constructs an n×n - dimensional perspective difference matrix M, where the matrix element M ij The calculation formula is:

[0091] where: t i and t j are the central moments of the i - th and j - th time intervals respectively; the time difference term in the denominator makes the influence of the perspective jump between intervals with larger time intervals smaller. Each element M ij in the matrix represents the severity of the perspective change when switching from the i - th time interval to the j - th time interval. The larger the value, the more obvious the perspective jump, and more careful handling of the switching effect is required. The perspective difference matrix of this implementation is shown in Table 1.

[0092] Table 1 Perspective Difference Matrix

[0093] T1 T2 T3 T4 T1 0 0.3 0.6 0.4 T2 0.3 0 0.8 0.5 T3 0.6 0.8 0 0.2 T4 0.4 0.5 0.2 0

[0094] S5. Use the video frame dataset to segment the first video stream to obtain the first video sequence, including: perform scene content analysis on the first video stream, extract scene change features according to the depth - of - field feature data and scene location data; according to the scene change features, use the threshold detection algorithm to identify the scene switching points in the first video stream, and divide the first video stream into multiple initial scene segments {SP1, SP2,......, SP l};

[0095] Calculate the content score value E i of each initial scene segment SP i , E i = λ1×C i + λ2×S i , Ci Represents the clarity score, HL j Represents the clarity heat map of the j-th frame in the segment SP i , where n i is the number of frames of the segment SP i , and max(HL j ) represents the maximum clarity value of the scene main area; S i Represents the stability score, Represents the variance of the clarity difference values between consecutive frames in the segment SP i ; for the scoring distribution of this embodiment, see Figure 9 .

[0096] Select segments with the content score value E i greater than the threshold as candidate segments to obtain the first video sequence SEQ1, SEQ1 = {SEQ 11 , SEQ 12 ,......, SEQ 1r}, and its corresponding time index information T1 = {T 11 , T 12 ,....., T 1r}.

[0097] S6. According to the first video sequence, use the video frame dataset and the perspective difference matrix to clip the second video stream to obtain the second video sequence as the clipped short video, including: The system retrieves the corresponding second video stream segment in the video frame dataset after time synchronization according to the time index information T1 of the first video sequence SEQ1. For example, for the segment SEQ 1,i in the first video sequence and its time index T 1,i = [t start , t end , the system searches for the frames in the second video stream whose time range overlaps with T 1,i to form the corresponding segment US 2,i . After completing the corresponding search for all segments, a segment set US2 = {US 21 , US 22 ,....., US 2r} of the second video stream is formed, as shown in Figure 10 . The blue is the first video stream, the purple is the second video stream, and the white frame is the corresponding selected segment. It should be noted that since the two video streams may not be completely synchronized, the system will adjust the time index according to the aforementioned time mapping function to ensure the accuracy of content correspondence.

[0098] Using the perspective difference matrix M, calculate the perspective complementarity VC between each scene segment in the first video sequence SEQ1 and the corresponding time period of the second video stream, VC = γ × Mij +(1 - γ)×(1 - S i ), where M ij is the corresponding element in the perspective difference matrix, representing the degree of difference between two perspectives; S i is the content similarity value for the corresponding time period, obtained by calculating in step S4; γ is the weight coefficient, usually set within the range of 0.3 to 0.7 and adjusted according to the editing style; this formula weighs the perspective difference (the first term) and the content complementarity (the second term).

[0099] When calculating the perspective complementarity degree, the system uses the sliding window technique to score the sub - segments in the second video stream US 2,i . The window size can vary from 0.5 seconds to several seconds, depending on the rhythm sense of the required editing style. The system sets the perspective complementarity degree threshold VC th (usually 0.5 - 0.7), and selects the segments with perspective complementarity degree VC > VC th to be included in the initial selection segment set US2'. In practical applications, the system will make judgments by combining the following factors: whether the segment duration meets the minimum duration requirement (usually ≥ 1 - 2 seconds); whether the content stability of the segment reaches the basic standard; whether the segment has content coherence with the previously selected segments.

[0100] The system arranges the segments in the initial selection segment set US2' in the original time order to form the initial version of the second video sequence SEQ2. To improve the video viewing experience, the system automatically selects appropriate transition effects according to the visual characteristics of adjacent segments: for adjacent segments with small perspective differences, a hard cut can be used; for adjacent segments with large perspective differences, a cross dissolve, fade - in and fade - out or other smooth transition effects can be used; the system dynamically determines the type and parameters of the transition effect (such as the dissolve duration) according to the magnitude of the corresponding value in the perspective difference matrix.

[0101] After the above - mentioned processing, the system outputs the final second video sequence SEQ2 as the short video finished product of the editing. While retaining the essence of the original content, this sequence provides a richer and smoother visual experience through the intelligent switching of multiple perspectives.

Claims

1. A short video editing method, characterized in that, Including: S1, obtaining multiple video streams associated with the target scene, as well as the position data and shooting parameters of each camera device; the multiple video streams include a first video stream and a second video stream, and the first video stream represents the main perspective video data as the editing benchmark; The second video stream represents the auxiliary perspective video data; S2, constructing a three-dimensional scene space model according to the position data and shooting parameters of each camera device; S3, extracting the depth of field feature data and scene position data of the multiple video streams; And performing time synchronization on the video frames of the multiple video streams to obtain a video frame data set; S4, calculating the perspective difference matrix of each camera device according to the depth of field feature data and scene position data; S5, performing segmentation processing on the first video stream by using the video frame data set to obtain a first video sequence; S6, according to the first video sequence, using the video frame data set and the perspective difference matrix to edit the second video stream to obtain a second video sequence as the edited short video.

2. The short video editing method according to claim 1, wherein: S2, constructing a three-dimensional scene space model, including: Extracting spatial coordinate data from the position data of each camera device, and extracting lens focal length, field of view angle and direction vector data according to the shooting parameters; Establishing a spatial position distribution model of the camera devices according to the spatial coordinate data, and calculating the shooting range projection data of each camera device according to the lens focal length, field of view angle and direction vector data; Constructing a three-dimensional scene space model of the target scene according to the spatial position distribution model and the shooting range projection data.

3. The short video editing method according to claim 2, wherein: S3, extracting the depth of field feature data and scene position data of the multiple video streams, including: Performing frame decomposition processing on the multiple video streams to generate an initial video frame set; Using a deep learning model to perform image analysis on the initial video frame set, identifying the clear area and blurred area of each video frame, and generating clarity heat map data; the clear area represents the focused area with clear imaging in the image, and the blurred area represents the non-focused area with blurred imaging in the image; Calculating the depth of field feature data of each video frame according to the clarity heat map data, and the depth of field feature data represents the focal depth distribution of different regions in the video frame; Performing object detection on the initial video frame set to identify the scene main body in the video frame, and the scene main body represents the main shooting object in the video; Mapping the scene main body to the three-dimensional scene space model, and calculating the position data of the scene main body in the three-dimensional scene space model as the scene position data.

4. The short video editing method according to claim 3, wherein: Obtaining the video frame data set, including: Extracting the timestamps of the multiple video streams, and establishing the initial time relationship of the multiple video streams according to the timestamps; According to the initial time relationship, combining the scene position data, matching the video frames of the same scene main body in different video streams; Performing time interpolation on the multiple video streams according to the matching result to obtain a time-calibrated video frame data set.

5. The short video editing method according to any one of claims 2 to 4, wherein: S4, calculating the perspective difference matrix of each camera device, including: Based on the depth-of-field feature data, compare and analyze the clarity heat map data of each video frame in the first video stream and the second video stream to generate video frame clarity difference data; Map the scene location data into a three-dimensional scene space model, and calculate the spatial shooting angle difference data of different camera devices for the same scene subject; According to the video frame clarity difference data and the spatial shooting angle difference data, calculate the content similarity data between the first video stream and the second video stream through a weighted fusion algorithm; According to the content similarity data, construct a perspective difference matrix reflecting the degree of perspective difference between each camera device.

6. The short video editing method according to claim 5, wherein: Generating video frame clarity difference data includes: Using the depth of field feature data, extract the sharpness heat map data of k pairs of video frames corresponding to the same time in the first video stream and the second video stream, and obtain the paired heat map data sets H1 = {H 11 , H 12 ,...., H 1k} and H2 = {H 21 , H 22 ,...., H 2k}; For the heatmap data (H 1j , H 2j ) of each pair of video frames, perform processing using bilinear interpolation or bicubic interpolation to obtain datasets H 1j ' and H 2j ' with the same resolution; According to the three-dimensional scene space model, using the internal and external parameters of the camera, construct the spatial mapping function M j ; Using function M j Map H 2j ' to the corresponding spatial position of H 1j ' to obtain H 2j ”; Calculate H 1j ' and H 2j ” the pixel - level difference value between |H 1j '(x, y) - H 2j ”(x, y)|, to obtain the difference matrix D 1j , where, H 1j '(x, y) represents the clarity value of the normalized heat map of the first video stream at the coordinate (x, y), and H 2j ”(x, y) represents the clarity value of the heat map of the mapped second video stream at the same coordinate; Calculate the average value μ of each difference matrix D 1j to obtain the single-frame sharpness difference value; j ​ The clarity difference value μ of all pairs of video frames j is used to form a difference data set μ = {μ1, μ2,....., μ k}, which serves as the video frame clarity difference data.

7. The short video editing method according to claim 5, wherein: Calculating the spatial shooting angle difference data of different camera devices for the same scene subject includes: Based on the scene location data obtained in step S3, determine the set of coordinate points P = {p1, p2,....., p n } in the three-dimensional scene space for the scene main body; Using the three-dimensional scene space model, extract the set of spatial position coordinates C1 = {C 11 , C 12 ,....., C 1m} and C2 = {C 21 , C 22 ,....., C 2m} of m pairs of camera devices corresponding to each frame in the first video stream and the second video stream; For each pair of camera device positions (C 1i , C 2i ), calculate the direction vectors V 1i and V 2i from positions C c and C 1i to the center point p 2i of the set of scene body coordinate points P; From the video frame dataset obtained in step S3, for each pair of camera device positions (C 1i , C 2i ), extract the image region containing the scene main body from the corresponding video frames; For each pair of camera device positions (C 1i , C 2i ), for the corresponding scene main body image area, use the SIFT or ORB algorithm to extract the feature point sets F 1i and F 2i ; For each pair of feature point sets (F 1i , F 2i ), calculate the feature point matching rate where N i,match is the number of successfully matched feature point pairs; simultaneously calculate the deformation matrix T i , where x1 and x2 are the homogeneous coordinates of the corresponding feature points in F 1i and F 2i respectively; Calculate the angle θ 1i between each pair of direction vectors V 2i and V i ; For each pair of camera device positions (C 1i , C 2i ), according to the feature point matching rate R i , the eigenvalues and included angle θ i of the deformation matrix T i , calculate the spatial shooting angle difference value D i : wherein, w1, w2, and w3 are weight coefficients and satisfy w1 + w2 + w3 = 1, σ i,max and σ i,min are respectively the maximum and minimum non-zero singular values of the deformation matrix T i ; The spatial shooting angle difference value D of all camera device pairs i constitute a difference data set D = {D1, D2,......, D m}, which serves as the spatial shooting angle difference data of different camera device pairs for the same scene subject.

8. The short video editing method according to claim 5, wherein: Calculating the content similarity data between the first video stream and the second video stream through a weighted fusion algorithm includes: Establish the temporal correspondence relationship between the video frame sharpness difference dataset μ = {μ1, μ2,......, μ k} and the spatial shooting angle difference dataset D = {D1, D2,......, D m}, and obtain the paired difference dataset (μ i ', D i '), i = 1, 2,...., q, where q is the number of pairs; For each pair of paired difference data (μ i ', D i '), calculate the content similarity value S i : S i = 1 - [α × norm(μ i ') + β × norm(D i '), where α and β are weight coefficients and satisfy α + β = 1, and norm(*) is a normalization function that maps the difference value to the interval [0, 1]; Take the content similarity value sequence {S1, S2,......, S q} as the content similarity data between the first video stream and the second video stream.

9. The short video editing method according to claim 5, wherein: Constructing a perspective difference matrix reflecting the degree of perspective difference between each camera device includes: Divide the sequence of content similarity values {S1, S2,......, S q} of the first video stream and the second video stream into n subsequences according to time periods, and each subsequence corresponds to a scene time interval; Calculate the statistical eigenvalue of each subsequence as the similarity feature vector SV for the corresponding time interval i ; According to the similarity feature vector SV i , calculate the viewing angle difference degree d between the first video stream and the second video stream in each time interval i , d i = g(SV i ), where the mapping function d i = g(SV i ) = 1 - (w1 × mean(SV i ) + w2 × (1 - var(SV i )) + w3 × (1 - max(SV i )) + w4 × min(SV i ))), where w1, w2, w3, and w4 are weighting coefficients and satisfy w1 + w2 + w3 + w4 = 1; mean(SV i ) represents the mean of SV i ; var(SV i ) represents the variance of SV i ; max(SV i ) represents the maximum value of SV i ; min(SV i ) represents the minimum value of SV i ; Construct an n×n - dimensional perspective difference matrix M, where the matrix element M ij represents the degree of perspective jump between the video content of the i - th time interval and the video content of the j - th time interval. The calculation formula is: where, t i and t j are the central moments of the i - th and j - th time intervals respectively.

10. The short video editing method according to claim 9, wherein: S5, obtaining the first video sequence, including: Perform scene content analysis on the first video stream, and extract scene change features according to the depth-of-field feature data and the scene location data; According to the scene change characteristics, use the threshold detection algorithm to identify the scene switching points in the first video stream, and divide the first video stream into multiple initial scene segments {SP1, SP2,......, SP l}; Calculate the content score value E of each initial scene segment SP i The content score value E i , E i = λ1×C i + λ2×S i , C i represents the clarity score, HL j represents the clarity heat map of the j-th frame in the segment SP i , n i is the number of frames of the segment SP i , max(HL j ) represents the maximum clarity value of the scene main area; S i represents the stability score, represents the variance of the clarity difference value between consecutive frames in the segment SP i ​ Selected content score value E i Fragments greater than the threshold are used as candidate fragments to obtain the first video sequence SEQ1, SEQ1 = {SEQ 11 , SEQ 12 ,......, SEQ 1r}, and its corresponding time index information T1, T1 = {T 11 , T 12 ,....., T 1r}.

Citation Information

Patent Citations

  • Automatic video editing method and device and electronic equipment

    CN119676510A