A video processing method and apparatus

CN122845859APending Publication Date: 2026-09-29HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611305203.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

在长视频匹配中,DTW算法为了迁就整体路径的平滑,导致局部核心画面(高光时刻(Highlight Time))出现几秒甚至更严重的偏移

Benefits of technology

[0021]经由上述技术方案可知,本申请在动态时间规整算法的基础上,引入高光区域锚点权重掩码构建“中间刚性约束、两端弹性适应”的锚点约束代价矩阵,锚点约束代价矩阵用于在源视频的高光时刻对应的行索引上施加奖励权重(降低代价),在远离高光时刻的区域施加惩罚权重(增加代价),在锚点约束代价矩阵上执行受限动态时间规整,受限动态时间规整在递推过程中强制去除水平方向和垂直方向的状态转移,要求时间轴单调向前推进以避免匹配至静止画面,从而在寻优阶段即消除传统算法的累积漂移问题。通过锚点约束代价场的锚点约束,确保商业价值最高的高光时刻误差控制在帧级范围内且不受片头片尾等剪辑影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845859A_ABST
    Figure CN122845859A_ABST
Patent Text Reader

Abstract

This application discloses a video processing method and apparatus, relating to the fields of multimedia data processing and computer vision technology. This solution introduces a highlight region anchor point weight mask to construct an anchor point constraint cost matrix. This matrix applies reward weights to the row indices corresponding to highlight moments in the source video and penalty weights to regions far from highlight moments. Constrained dynamic time warping is performed on the anchor point constraint cost matrix. During the recursive process, constrained dynamic time warping forcibly removes horizontal and vertical state transitions, requiring the time axis to monotonically advance to avoid matching static images, thus eliminating the cumulative drift problem of traditional algorithms during the optimization phase. Through the anchor point constraint cost, the error of the highlight moments with the highest commercial value is ensured to be controlled within the frame-level range, unaffected by editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of multimedia data processing and computer vision technology, and more specifically, to a video processing method and apparatus. Background Technology

[0002] In video content security, copyright protection, and advertising monitoring, video matching is used to determine whether segments of the source video (the unprocessed initial video) appear in the target video (the source video after editing, inserting advertisements, etc.).

[0003] Existing video matching technologies, such as sliding window or hash fingerprint matching, assume that the video is played at a constant speed. Once the target video undergoes non-linear editing (such as slow-motion close-ups, fast-forward transitions, speed changes, etc.) and visual manipulation (such as filters, flipping, picture-in-picture, etc.), linear matching will completely fail, causing the linear matching to fail. Furthermore, existing video matching technologies can handle temporal scaling through Dynamic Time Warping (DTW) algorithms, but general DTW algorithms aim to minimize the global path cost. In long video matching, in order to accommodate the smoothness of the overall path, DTW algorithms cause local core frames (highlight moments) to be shifted by several seconds or even more.

[0004] Therefore, how to avoid matching failures when faced with non-linear editing and visual secondary creation interference, and how to avoid offsets during video matching, are the problems that this application urgently needs to solve. Summary of the Invention

[0005] In view of this, this application discloses a video processing method and apparatus, which aims to avoid matching failures when faced with non-linear editing and visual secondary creation interference, and to avoid offsets during video matching.

[0006] To achieve the above objectives, the disclosed technical solution is as follows:

[0007] The first aspect of this application discloses a video processing method, the method comprising:

[0008] Obtain the first frame sequence of the source video and the second frame sequence of the target video;

[0009] Semantic features are extracted from the first frame sequence to obtain the first semantic feature vector sequence of the source video, and semantic features are extracted from the second frame sequence to obtain the second semantic feature vector sequence of the target video.

[0010] A basic distance matrix is ​​constructed using the first semantic feature vector sequence and the second semantic feature vector sequence;

[0011] The anchor point constraint cost matrix is ​​determined based on the base distance matrix and the predefined anchor point weight function;

[0012] The optimal path is determined based on the anchor point constraint cost matrix.

[0013] The optimal path is analyzed to obtain the analysis results, and the playback strategy of the target video is deduced in reverse based on the analysis results.

[0014] A second aspect of this application discloses a video processing apparatus, the apparatus comprising:

[0015] The acquisition unit is used to acquire the first frame sequence of the source video and the second frame sequence of the target video.

[0016] The extraction unit is used to extract semantic features from the first frame sequence to obtain the first semantic feature vector sequence of the source video, and to extract semantic features from the second frame sequence to obtain the second semantic feature vector sequence of the target video.

[0017] The construction unit is used to construct a basic distance matrix using the first semantic feature vector sequence and the second semantic feature vector sequence;

[0018] The first determining unit is used to determine the anchor point constraint cost matrix based on the base distance matrix and the predefined anchor point weight function.

[0019] The second determining unit is used to determine the optimal path based on the anchor point constraint cost matrix.

[0020] The analysis and derivation unit is used to analyze the optimal path to obtain the analysis result, and to deduce the playback strategy of the target video in reverse based on the analysis result.

[0021] As can be seen from the above technical solution, this application, based on the dynamic time warping algorithm, introduces a highlight region anchor point weight mask to construct an anchor point constraint cost matrix with "rigid constraints in the middle and elastic adaptation at both ends". The anchor point constraint cost matrix is ​​used to apply reward weights (reduce costs) to the row indices corresponding to the highlight moments in the source video, and to apply penalty weights (increase costs) to regions far from the highlight moments. Restricted dynamic time warping is performed on the anchor point constraint cost matrix. During the recursive process, restricted dynamic time warping forcibly removes the state transitions in the horizontal and vertical directions, requiring the time axis to advance monotonically to avoid matching to static images, thereby eliminating the cumulative drift problem of traditional algorithms in the optimization stage. Through the anchor point constraints of the anchor point constraint cost field, it is ensured that the error of the highlight moments with the highest commercial value is controlled within the frame-level range and is not affected by editing such as intros and outros. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating a video processing method disclosed in an embodiment of this application;

[0024] Figure 2 This is a schematic diagram illustrating the principle of dual-stream semantic feature extraction and anchor point constraint cost matrix construction disclosed in the embodiments of this application;

[0025] Figure 3 This is a schematic diagram of the path slope analysis and playback strategy reverse logic disclosed in the embodiments of this application;

[0026] Figure 4 This is a schematic diagram of the structure of a video processing apparatus disclosed in an embodiment of this application;

[0027] Figure 5 This is a schematic diagram of the structure of the electronic device disclosed in the embodiments of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0030] As the background technology indicates, existing video matching techniques, such as sliding window or hash fingerprint matching, assume that the video is played at a constant speed. Once the target video undergoes non-linear editing and visual manipulation, linear matching will completely fail, causing the linear matching to fail. Furthermore, existing video matching techniques can handle temporal scaling through the DTW algorithm, but the general DTW algorithm aims to minimize the global path cost. In long video matching, the DTW algorithm, in order to accommodate the smoothness of the overall path, causes local core frames to shift.

[0031] To address the aforementioned issues, this application discloses a video processing method and apparatus. Based on the dynamic time warping algorithm, this solution introduces a highlight region anchor point weight mask to construct an anchor point constraint cost matrix with "rigid constraints in the middle and flexible adaptation at both ends." This anchor point constraint cost matrix applies reward weights (reducing costs) to the row indices corresponding to the highlight moments in the source video and penalty weights (increasing costs) to regions far from the highlight moments. Restricted dynamic time warping is then performed on the anchor point constraint cost matrix. During the recursive process, restricted dynamic time warping forcibly removes horizontal and vertical state transitions, requiring the time axis to monotonically advance to avoid matching to static images, thus eliminating the cumulative drift problem of traditional algorithms during the optimization phase. Through the anchor point constraints of the anchor point constraint cost field, the error of the commercially valuable highlight moments is ensured to be controlled within the frame-level range and is unaffected by editing such as intros and outros, while retaining the ability to flexibly match non-highlight regions, avoiding matching failures when facing non-linear editing and visual secondary creation interference. Furthermore, semantic features are extracted from both the first and second frame sequences using a multi-stream semantic feature extraction method to enhance the recognition of the main content and resist background noise. Local slope analysis is performed on the optimal path output by the constrained dynamic time warping, and the playback strategy of the target video is derived inversely based on the local slope. The specific implementation is illustrated in the following embodiments.

[0032] It should be noted that the video processing method and apparatus provided in this application can be used in the fields of multimedia data processing and computer vision, specifically involving a method for achieving high-precision temporal positioning, flexible mapping, and playback strategy analysis of short video clips in long videos under complex nonlinear editing scenarios. The above is only an example and does not limit the application field of the video processing method and apparatus provided in this application.

[0033] refer to Figure 1 The image shows a video processing method disclosed in an embodiment of this application. The video processing method mainly includes the following steps:

[0034] S101: Obtain the first frame sequence of the source video and the second frame sequence of the target video.

[0035] The source video is the original video that has not been edited or had any ads inserted.

[0036] The target video is the video after the source video has been edited, and advertisements have been inserted.

[0037] A multi-level coarse-to-fine search is employed to obtain the first frame sequence of the source video and the second frame sequence of the target video. Specifically, a multi-level search strategy is used to first perform sparse sampling of the video at 1fps, then use low-dimensional features to perform fast DTW to lock several candidate time windows, and then restore a high frame rate within the candidate time windows for fine matching to obtain the first frame sequence of the source video and the second frame sequence of the target video.

[0038] To address the computational performance bottleneck in target video matching (such as a 2-hour movie), this solution employs a multi-level search strategy, which includes a first-level coarse screening (Global Coarse Filter) and a second-level fine search to obtain the first frame sequence of the source video and the second frame sequence of the target video.

[0039] The specific process of obtaining the first frame sequence of the source video and the second frame sequence of the target video is shown in A1-A4.

[0040] A1: Perform sparse sampling on the acquired source and target videos.

[0041] The A1 process is a first-level coarse screening downsampling process. Specifically, the source video and the target video are sparsely sampled according to frames per second, such as 1 fps.

[0042] A2: Execute DTW to filter out all matching segments.

[0043] In A2, perform wide threshold recall, i.e., execute fast DTW matching, such as a low similarity threshold, with a similarity set to 0.4, to filter out all possible matching segments.

[0044] A3: Perform operations on all matching segments to form candidate time windows for the target video.

[0045] In A3, window redundancy is applied to all matching segments, that is, all matching segments are extended forward and backward by 30 seconds (padding) to form candidate time windows, ensuring that the true value is not missed.

[0046] Operations include, but are not limited to, extended operations.

[0047] The process from A1 to A3 is the first-stage coarse screening process.

[0048] A4: Perform frame rate recovery sampling within the candidate time window to obtain the source video after frame rate recovery and normalize it to obtain the first frame sequence, and obtain the second frame sequence of the target video after frame rate recovery and normalize it to obtain the second frame sequence.

[0049] In A4, within the candidate time window, the preset frame rate is restored, i.e., high frame rate sampling (8fps-10fps recommended) is performed within the candidate time window, and the image is normalized to, for example, 224×224 pixels for subsequent fine calculations. This scheme uses multi-level cascaded search to prevent missed detections and optimize performance. Normalization is used to reduce computational load and improve efficiency.

[0050] The A4 process is the preparation process for the secondary detailed search.

[0051] In this embodiment, a multi-level cascaded search strategy (first sparse sampling, then fast DTW to filter matching segments, then expanding to form a candidate window, and finally restoring and normalizing the frame rate within the window) is adopted to first perform coarse screening and then fine search, which effectively solves the computational performance bottleneck of long video matching, prevents missed detections, reduces computational load, and improves overall processing efficiency.

[0052] S102: Extract semantic features from the first frame sequence to obtain the first semantic feature vector sequence of the source video, and extract semantic features from the second frame sequence to obtain the second semantic feature vector sequence of the target video.

[0053] The first semantic feature vector sequence includes the first visual feature vector sequence and the first spectral feature vector sequence of the source video.

[0054] The second semantic feature vector sequence includes the second visual feature vector sequence of the target video and the second spectral feature vector sequence of the target video.

[0055] The specific process of extracting semantic features from the first frame sequence to obtain the first semantic feature vector sequence of the source video is shown in B1-B3.

[0056] B1: Extract the first global semantic vector and the first subject semantic vector from each frame of the first frame sequence.

[0057] B2: The first global semantic vector and the first subject semantic vector are concatenated to generate the first visual feature vector sequence of the source video.

[0058] For the first frame sequence, the following feature encoding process is performed:

[0059] Visual feature extraction (VisualEmbedding): Pre-trained multimodal models such as CLIP (ViT-B / 32) can be used.

[0060] Full-image stream: Input the normalized first frame sequence, extract the first global semantic vector, and then... express.

[0061] Central Flow: A central region (e.g., 30%-80%) encompassing the main content is cropped out and input into the model to extract the first subject semantic vector. This first subject semantic vector is used to combat picture-in-picture interference. The first subject semantic vector is obtained through... express.

[0062] The first global semantic vector and the first subject semantic vector are concatenated to generate the first visual feature vector sequence of the source video. The feature concatenation formula is shown in formula (1):

[0063] (1);

[0064] in, The first visual feature vector sequence of the source video; This is the first global semantic vector; This is the first subject semantic vector; For splicing.

[0065] Using multimodal models such as CLIP, the full-image features (i.e., the first global semantic vector) and the central region features (i.e., the first subject semantic vector) of the normalized first frame sequence are extracted and concatenated to enhance the recognition of the subject content and resist background noise (such as subtitles, bullet comments, and picture-in-picture borders).

[0066] B3: Extract audio features from the audio stream corresponding to the first frame sequence to obtain the first spectral feature vector sequence of the source video.

[0067] In B3, audio feature extraction (AudioEmbedding) is performed on the corresponding audio stream from the first frame sequence to extract the first spectral (Log-Mel) feature vector sequence of the corresponding audio segment from the source video. This first spectral feature vector sequence is then processed... express.

[0068] In this embodiment, global semantic vectors and main semantic vectors are extracted from the frame sequence and concatenated, while audio spectral features are also extracted. Multi-stream semantic feature fusion enhances the ability to recognize the main content of the video, effectively resists complex visual interference (such as picture-in-picture, subtitles, bullet comments, filters, etc.), and improves the robustness of matching.

[0069] The specific process of extracting semantic features from the second frame sequence to obtain the second semantic feature vector sequence of the target video is shown in C1-C3.

[0070] C1: Extract the second global semantic vector and the second subject semantic vector from each frame of the second frame sequence.

[0071] C2: The second global semantic vector and the second subject semantic vector are concatenated to obtain the second visual feature vector sequence of the target video.

[0072] For the second frame sequence in the refinement phase, the following feature encoding process is performed:

[0073] Visual feature extraction (VisualEmbedding): Pre-trained models such as CLIP (ViT-B / 32) can be used.

[0074] Full-image stream: Input the normalized second frame sequence, extract the second global semantic vector, and then... express.

[0075] Central Flow: A central region (e.g., 30%-80%) encompassing the main content is cropped and input into the model to extract a second main semantic vector. This second main semantic vector is used to combat picture-in-picture interference. The second main semantic vector is obtained through... express.

[0076] The second global semantic vector and the second main semantic vector are concatenated to generate the second visual feature vector sequence of the target video. The feature concatenation formula is shown in formula (2):

[0077] (2);

[0078] in, The second visual feature vector sequence of the target video; This is the second global semantic vector of the target video; This is the second main semantic vector of the target video.

[0079] Using multimodal models such as CLIP, the full-image features (i.e., the second global semantic vector) and central region features (i.e., the second main semantic vector) of the normalized second frame sequence are extracted and concatenated to enhance the recognition of the main content and resist background noise (such as subtitles, bullet comments, and picture-in-picture borders).

[0080] C3: Extract audio features from the corresponding audio stream of the second frame sequence to obtain the second spectral feature vector sequence of the target video.

[0081] In C3, audio features are extracted from the corresponding audio stream in the second frame sequence to obtain the second spectral (Log-Mel) feature vector sequence of the corresponding audio segment of the target video. The second spectral feature vector sequence is then processed... express.

[0082] The standard deviation of elements in the first semantic feature vector sequence and / or the second semantic feature vector sequence is calculated respectively. This standard deviation is then compared with a preset threshold. If the standard deviation is less than the preset threshold (the threshold is set according to actual conditions, and this application does not specify a particular threshold), the corresponding image frame is determined and marked as an invalid frame with no information content. When calculating the basic distance between the source video and the target video, the visual and audio weights of the invalid frame are adaptively adjusted. By calculating the standard deviation of the semantic feature vector sequence and comparing it with a threshold, invalid frames (such as black screens or solid colors) are determined and marked. The adaptive adjustment of visual and audio weights during distance calculation has the beneficial effect of achieving dynamic adaptive weighting of audio and video based on image information entropy. When the image has no information content, it automatically relies on audio features, effectively resisting interference from meaningless frames and further enhancing the robustness of the matching.

[0083] Specifically, when calculating the basic distance between the source video and the target video, the visual weight of invalid frames is reduced according to the marking, and the corresponding audio weight is increased accordingly, so as to complete the dynamic adaptive adjustment of audio and video weights based on image information entropy, i.e., dynamic weight decision.

[0084] By employing multi-stream semantic feature fusion and information entropy-based dynamic audio-visual weights, complex visual interference can be resolved.

[0085] S103: Construct the basic distance matrix using the first semantic feature vector sequence and the second semantic feature vector sequence.

[0086] The specific execution process of S103 is shown in D1-D3.

[0087] D1: Calculate the cosine distance between each feature vector in the first semantic feature vector sequence and each feature vector in the second semantic feature vector sequence.

[0088] D2: Get the length of the first frame sequence of the source video and the length of the second frame sequence of the target video.

[0089] Let the length of the first frame sequence of the source video be M, and the length of the second frame sequence of the target video be N. The specified highlight moment frame index is... .

[0090] D3: Calculate the fundamental distance matrix between the source video and the target video based on the cosine distance, the length of the first frame sequence of the source video, and the length of the second frame sequence of the target video.

[0091] The basic distance matrix is ​​calculated as shown in formula (3).

[0092] (3);

[0093] in, Use the basic distance matrix.

[0094] For each node of the base distance matrix, calculate the cosine distance between the first semantic feature vector sequence and the second semantic feature vector sequence. The formula for calculating the cosine distance is shown in formula (4).

[0095] (4);

[0096] in, The cosine distance between each feature vector in the first semantic feature vector sequence and each feature vector in the second semantic feature vector sequence (i.e., the combined matching cost between the i-th frame and the j-th frame). These are preset visual feature weights, used to control the proportion of image features in the overall distance; It is the first visual feature vector of the i-th frame in the source video (each frame feature vector in the first semantic feature vector sequence includes the first visual feature vector of the corresponding frame in the source video). The second visual feature vector of the j-th frame in the candidate time window of the target video (each frame feature vector in the second semantic feature vector sequence includes the second visual feature vector of the corresponding frame of the source video). These are preset audio feature weights, used to control the proportion of audio features in the overall distance; Let be the first audio spectrum feature vector corresponding to the i-th frame in the source video; The second audio spectral feature vector corresponding to the j-th frame in the candidate time window of the target video.

[0097] If the image frame marked above is an invalid frame, then , (Fully dependent on audio).

[0098] If the image frame marked above is a valid frame, i.e., its standard deviation is greater than or equal to the threshold, then the frame corresponding to that information entropy is marked as a valid frame. , (primarily visual).

[0099] In this embodiment, the cosine distance between each feature vector in the first semantic feature vector sequence and each feature vector in the second semantic feature vector sequence is calculated, and a basic distance matrix is ​​constructed by combining the frame sequence length. Based on the comprehensive distance metric of multimodal semantic features (visual + audio), it can accurately reflect the similarity between the source video and the target video frame, and provide a reliable basic matching for subsequent anchor point constraints and path planning.

[0100] S104: Determine the anchor point constraint cost matrix based on the base distance matrix and the predefined anchor point weight function.

[0101] In S104, construct the anchor weight mask. This involves constructing a row vector to adjust the matching cost of different rows, defining an anchor weight function, and using the anchor weight function through... express.

[0102] The specific process for determining the anchor point constraint cost matrix is ​​as follows: An anchor point weight mask based on a flat-top Gaussian distribution is superimposed on the calculated base distance matrix. This anchor point weight mask includes reward weights and penalty weights. Reward weights are applied at preset positions corresponding to the target time (i.e., highlight moments) in the source video, and penalty weights are applied to regions corresponding to non-target times in the source video, resulting in the anchor point constraint cost matrix. The preset positions are the row indices corresponding to the target times in the source video. In this scheme, the row index refers to the matrix row coordinates of each frame of the source video in the base distance matrix and the cost matrix.

[0103] Core adsorption region ( ): The weight coefficient of the core adsorption region is a positive number less than 1.0. In this embodiment, the preferred range of the weight coefficient of the core adsorption region is 0.3 to 0.7. The smaller the value of the weight coefficient of the core adsorption region, the stronger the adsorption ability of keyframes; the larger the value of the weight coefficient of the core adsorption region, the higher the sensitivity to feature differences.

[0104] Non-core area: .

[0105] A coefficient greater than 1 is set as the "penalty," and the further away from the highlight anchor point, the heavier the penalty. The recommended value is 5.0.

[0106] in, Anchor weight function; The penalty intensity is represented by exp; exp is the exponentiation operation. To control bandwidth; This is the frame index for the highlight moments.

[0107] In the embodiments of this application, an anchor weight mask based on a flat-top Gaussian distribution is superimposed on the calculated basic similarity matrix. A "reward weight" (reducing the cost) is applied to the row index (i.e., the frame index of the highlight moment) corresponding to the highlight moment in the source video, while a "penalty weight" (increasing the cost) is applied to regions far from the highlight moment. and Together, they constructed a non-uniform "funnel-shaped" anchor point constrained cost field.

[0108] The value of (recommended 5.0) determines the edge potential energy height of the confined field. A higher value... The value can effectively suppress the cumulative error (drift) of the DTW algorithm in sequence matching of the target video and prevent the path from deviating significantly in non-critical areas.

[0109] The value of determines the effective range of the constraint field. A reasonable This setting allows the algorithm to maintain local path flexibility near keyframes, thus ensuring compatibility with video compression, frame drops, or slight time shifts from editing. This can be achieved by adjusting... and By adjusting the ratio, this scheme achieves a dynamic balance between "rigid anchoring" and "elastic mapping".

[0110] The anchor point constraint cost matrix is ​​calculated as shown in formula (5).

[0111] (5);

[0112] in, For each element in the i-th row and j-th column of the anchor point constraint cost matrix; Basic distance matrix The element value corresponding to the i-th row and j-th column; Let be the weight value of the anchor weight function corresponding to the i-th row.

[0113] The specific principle for constructing the anchor point constraint cost matrix is ​​as follows: Figure 2 As shown. Figure 2 The logic of weighting the base distance matrix using a Gaussian weight mask is demonstrated.

[0114] Figure 2 In the input, the first frame sequence of the normalized source video and the second frame sequence of the normalized target video are used as inputs;

[0115] Each frame of the first frame sequence is input into the CLIP encoder to extract the first global semantic vector;

[0116] For each frame of the first frame sequence, the center is cropped by 50%, and the cropped center region image is input into the CLIP encoder to extract the first main semantic vector.

[0117] The first global semantic vector and the first main semantic vector are concatenated to obtain the first visual feature vector sequence of the source video;

[0118] Audio features are extracted from the audio stream corresponding to the first frame sequence to obtain the first spectral feature vector sequence of the source video;

[0119] The second frame sequence is input into the CLIP encoder to extract the second global semantic vector;

[0120] For each frame of the second frame sequence, the center is cropped by 50%, and the cropped center region image is input into the CLIP encoder to extract the second main semantic vector.

[0121] The second global semantic vector and the second main semantic vector are concatenated to obtain the second visual feature vector sequence of the target video.

[0122] Audio features are extracted from the audio stream corresponding to the second frame sequence to obtain the second spectral feature vector sequence of the target video;

[0123] Calculate the cosine distance between each feature vector in the first semantic feature vector sequence (first visual feature vector sequence and first spectral feature vector sequence) and each feature vector in the second semantic feature vector sequence (second visual feature vector sequence and second spectral feature vector sequence);

[0124] Obtain the length of the first frame sequence of the source video and the length of the second frame sequence of the target video;

[0125] Construct a basic distance matrix between the source video and the target video based on the cosine distance, the length of the first frame sequence, and the length of the second frame sequence;

[0126] A "reward weight (cost reduction)" is applied to the highlight frame index corresponding to the highlight moment in the source video. Based on the calculated basic similarity matrix, an anchor weight mask based on a flat-top Gaussian distribution is superimposed to obtain the anchor weight function. );

[0127] If in the core area ( ), The weight coefficient of the core area is a positive number less than 1.0, generating the anchor point weight function, through... express;

[0128] If in a non-core area (such as an edge area), set a coefficient greater than 1.0 as a penalty; if the weight coefficient is greater than 1.0, generate an anchor weight function.

[0129] The base distance matrix and anchor weight function are weighted row-wise. That is, the reward weight of the anchor weight mask is applied to the preset position corresponding to the target time in the source video, and the penalty weight of the anchor weight mask is applied to the region corresponding to the non-target time in the source video. The anchor constraint cost matrix is ​​determined, thereby achieving the effect of high-light rows forming low potential energy trenches that force the path to pass through the anchor points.

[0130] S105: Determine the optimal path based on the anchor point constraint cost matrix.

[0131] In S105, the first row of the anchor point constraint cost matrix is ​​initialized with cumulative cost. Based on the preset slope constraint condition, the state transition calculation is performed on the initialized anchor point constraint cost matrix to obtain the cumulative cost matrix. The minimum point in the last row of the cumulative cost matrix is ​​found as the endpoint. Starting from the endpoint, the reverse path backtracking is performed in the cumulative cost matrix to obtain the optimal path.

[0132] Constrained DTW is performed on the anchor point constraint cost matrix. The DTW algorithm with local slope constraints is executed to find the optimal path. Horizontal and vertical state transitions are forcibly removed, and the time axis is required to move monotonically forward to avoid matching static scenes.

[0133] 1. Initialization: Initialize the cumulative cost of all columns in the first row of the target video to C(0, j) (allowing entry from any position in the target video).

[0134] 2. State transition (with slope constraint), as shown in formula (6):

[0135] (6);

[0136] It should be noted that the states (i, j-1) and (i-1, j) were deliberately excluded to force the timeline to flow monotonically forward, thus avoiding matching PPT slides or static images.

[0137] in, For each element in the cumulative cost matrix, there exists a value representing the cumulative cost of moving from the starting point to the current point (i, j). These are elements in the anchor point constraint cost matrix, used to represent the current matching cost, indicating how similar the i-th frame of the source video and the j-th frame of the target video are; the more similar they are... The smaller the value (lower the cost), the less like The larger the value, the higher the cost.

[0138] Path backtracking involves finding the minimum value in the last line of the target video as the endpoint, and then using path backtracking to obtain the optimal path. The optimal path is then used to... express.

[0139] The optimal path consists of a series of coordinate points, and its expression is shown in formula (7):

[0140] (7);

[0141] Where x corresponds to the source video index i; y corresponds to the target video index j.

[0142] By initializing the anchor point constraint cost matrix, performing state transitions with slope constraints, finding the minimum value in the last row as the endpoint, and backtracking in reverse, the beneficial effects are that the horizontal and vertical state transitions are forcibly removed during the restricted DTW recursion process, avoiding matching static images, eliminating cumulative drift in long video matching, and ensuring the accuracy and reliability of the optimal path.

[0143] S106: Analyze the optimal path to obtain the analysis results, and deduce the playback strategy of the target video in reverse based on the analysis results.

[0144] The analysis results are used to represent the local slope of the path. In this scheme, the local slope is specifically applied to perform segmented analysis on the normalized optimal matching path. By using the local slope analysis of the path, the playback strategy (speed-up factor) of the target video can be derived in reverse.

[0145] The local slope includes a first slope and a second slope. The local slope is determined by... Indicates. The first slope passes through Indicated. The second slope passes through express.

[0146] The specific process involves analyzing the optimal path to obtain the local slope, and then using the local slope to deduce the playback strategy for the target video, as shown in E1-E4.

[0147] E1: Determine the cutting point.

[0148] In E1, find the highlight anchor point (i.e., the source index is...) in the optimal path. That point), denoted as .Will The corresponding target coordinate y is converted into a timestamp, which is the precise time of the target highlight.

[0149] The cutting points include the start point, the highlight anchor point, and the end point.

[0150] starting point: ;

[0151] Highlight anchor points: (Right now );

[0152] end: .

[0153] E2: Divide the optimal path by the cutting point to obtain the first and second path segments.

[0154] E3: Calculate the first slope of the first path segment and the second slope of the second path segment.

[0155] The first slope of the first path segment is calculated as shown in formula (8).

[0156] (8);

[0157] The second slope of the second path segment is calculated as shown in formula (9).

[0158] (9);

[0159] E4: Based on the first slope, the second slope, and the preset slope value, segmented speed analysis is performed to obtain the speed ratio of the first path segment and the speed ratio of the second path segment, thus completing the process of obtaining the playback strategy of the target video.

[0160] The preset slope value can be set to 0.5, 1.0, 2.0, etc. The preset slope value in this application is not specifically limited.

[0161] In this embodiment, the optimal path is divided into two segments by the cutting point, the first slope and the second slope are calculated respectively, and the segmented speed analysis is performed in combination with the preset slope value to obtain the speed ratio of each segment. This allows for accurate reverse derivation of the specific playback strategy of the target video at different time periods (such as original speed, slow motion, fast forward).

[0162] Segmented speed change analysis (track detective):

[0163] If k≈1.0, play at the original speed; it should be noted that in actual video processing and frame dropping situations, the calculated slope is basically an approximation of 0.98 or 1.02, so using the approximate equality sign is more in line with actual engineering logic.

[0164] If k≈0.5 (the source video moves 1 step and the target video moves 2 steps), it is judged as slow motion;

[0165] If k≈2.0 (the source video moves 2 steps and the target video moves 1 step), it is determined to be fast forward.

[0166] The specific analysis of segmented speed change is shown in Table 1.

[0167] Table 1

[0168]

[0169] The playback strategy for the target video is determined based on segmented speed analysis. Specifically, as follows... Figure 3 As shown. Figure 3 It provides path slope analysis and playback strategy inverse logic, and displays the physical playback state (fast forward / slow motion) corresponding to different slope values.

[0170] Figure 3 In the process, find the highlight anchor point in the optimal path (P); where the optimal path includes the source video index i and the target video index j;

[0171] The optimal path is divided by the highlight anchor point to obtain the first path (i.e. the first half of the path) and the second path (i.e. the second half of the path).

[0172] Calculate the first slope of the first path segment ( ), and calculate the second slope of the second path segment ( );

[0173] like ≈1.0, the first half plays at the original speed;

[0174] like =0.5, slow motion / lengthening of the first half;

[0175] like >1.0, such as 2.0, fast forward / compress the first half;

[0176] like ≈1.0, the second half plays at the original speed;

[0177] like ≈0.5, slow motion in the second half;

[0178] like >1.0, fast forward the second half;

[0179] Based on the first half being slow-motion / lengthened and the second half being played at original speed, output the playback strategy of the target video, namely, the first half being slow-motion + the second half being played at original speed.

[0180] Based on the fast-forward / compression of the first half and the slow motion of the second half, the playback strategy of the target video is output, namely fast-forward of the first half and slow motion of the second half.

[0181] The final output is shown in the example below:

[0182] {

[0183] "match_result": true,

[0184] "source_highlight": "00:00:15",

[0185] "target_highlight": "00:45:42", / / The precise time after mapping;

[0186] "edit_strategy": "slow_motion_start_fast_end" / / Speed ​​strategy analysis;

[0187] }

[0188] Closed-loop adaptive feedback is implemented, meaning adaptive parameter adjustments are made. Specifically, the system automatically records matching errors. If the highlight timing is manually corrected, the system automatically triggers a parameter update: increasing... and reduce This makes the constraints on the anchor point tighter in the next match.

[0189] By analyzing the local slope of the optimal path, the playback speed of the target video is inferred. Simultaneously, based on human-reported specular errors, the parameters of the weighted mask (penalty strength and bandwidth) are adaptively adjusted.

[0190] To facilitate understanding of the video processing process, an example is provided here. This example illustrates the overall workflow of flexible temporal mapping of video based on keyframe anchor point constraints.

[0191] Sparsely sample the source and target videos at 1fps;

[0192] Perform fast DTW matching, set a low similarity threshold (e.g., 0.4), and filter out all possible matching segments;

[0193] Padding the matching segment forward and backward by 30 seconds to lock the candidate time window;

[0194] Within the candidate time window, restore the high frame rate of 10fps and normalize the first frame sequence of the source video and the second frame sequence of the target video.

[0195] The first frame sequence after normalization is visually encoded using a pre-trained CLIP model to obtain the first global semantic vector and the first subject semantic vector. The second frame sequence is visually encoded using the pre-trained CLIP model to obtain the second global semantic vector and the second subject semantic vector.

[0196] The first global semantic vector and the first main semantic vector are concatenated to generate the first visual feature vector sequence of the source video;

[0197] Audio features are extracted from the audio stream corresponding to the first frame sequence to obtain the first spectral feature vector sequence of the source video;

[0198] The second global semantic vector and the second main semantic vector are concatenated to obtain the second visual feature vector sequence of the target video.

[0199] Audio features are extracted from the audio stream corresponding to the second frame sequence to obtain the second spectral feature vector sequence of the target video;

[0200] Dynamic weighting is applied to the first visual feature vector and the second visual feature vector to detect image information entropy, which involves calculating the element standard deviation (or the corresponding image frame information entropy) of the first visual feature vector itself or the second visual feature vector itself.

[0201] If the standard deviation (or information entropy) of an element is less than the threshold (the threshold is set according to the actual situation, and this application does not make a specific limitation), such as a black screen or a solid color, then the frame corresponding to the information entropy is marked as an invalid frame, and the weight is determined based on audio, that is, the weight is the high audio weight.

[0202] If the standard deviation (or information entropy) is greater than or equal to the threshold, the frame is marked as a valid frame, and the weight is high visual weight, which is mainly visual.

[0203] Based on the above formula (4), high audio weight or high visual weight, calculate the cosine distance between the first semantic feature vector sequence and the second semantic feature vector sequence;

[0204] Based on the base distance matrix and the predefined anchor weight function, the anchor constraint cost matrix is ​​determined. The construction of the anchor constraint cost matrix involves parameters such as penalty intensity, control bandwidth, core region, highlight moment frame index, and anchor weight mask.

[0205] Perform a restricted DTW path search from the anchor point constraint cost matrix to backtrack the optimal path;

[0206] The optimal path is mapped for highlight moments and the local slope of the path segments is analyzed to output the playback strategy of the target video (i.e., precise time + variable speed strategy).

[0207] The playback strategy for the target video is manually reviewed and feedback is provided to determine whether the error is greater than the preset difference (e.g., 0.5 seconds).

[0208] If the error is less than or equal to 0.5 seconds, the playback strategy for the target video passes the review.

[0209] If the error is greater than 0.5s, the playback strategy for the target video has not passed the review and adaptive parameter adjustment will be performed.

[0210] During the adaptive parameter adjustment process, it can be increased / Decrease This makes the constraints on the anchor point tighter in the next match.

[0211] After analyzing the optimal path to obtain the local slope and inversely deriving the playback strategy of the target video based on the local slope, the penalty intensity and bandwidth parameters of the flat-top Gaussian distribution are adaptively updated according to the closed-loop adaptive feedback mechanism. This automatically optimizes the anchor point constraint parameters, making the subsequent matching constraints on keyframes more accurate.

[0212] This solution aims to address the problem of existing video matching technologies failing to match when faced with non-linear editing (variable speed, fast forward, slow motion) and visual manipulation (filters, flipping, picture-in-picture). The goal is to achieve flexible mapping of short video clips into long videos and, through a unique "anchor point constraint" mechanism, to enforce absolute alignment accuracy of commercially valuable highlights on the timeline, preventing matching drift caused by accumulated errors in long videos.

[0213] This solution can replace manual review and verification of complex secondary creation videos, and is expected to reduce manual review costs by at least 80%. In advertising monitoring scenarios, it can accurately locate the highlights of ad display, avoid billing disputes caused by fast forward / slow motion, and has extremely high commercial monetization value.

[0214] The advantages of this solution are as follows:

[0215] 1) Absolute alignment accuracy: Through anchor point constraints, ensure that the error of the most commercially valuable highlight moments is controlled within the frame-level range, unaffected by editing such as intros and outros;

[0216] 2) Extremely strong anti-interference: semantic features can resist 90-degree flip, sketch filter, and 50% area occlusion;

[0217] 3) Intelligent strategy analysis: This system goes beyond simple "match / not match" binary judgment and can output specific speed change strategies (such as "2x speed in the first half and 0.5x speed in the second half").

[0218] In this embodiment, based on the dynamic time warping algorithm, a highlight region anchor point weight mask is introduced to construct an anchor point constraint cost matrix with "rigid constraints in the middle and flexible adaptation at both ends." This anchor point constraint cost matrix applies reward weights (reducing costs) to the row indices corresponding to the highlight moments in the source video and penalty weights (increasing costs) to regions far from the highlight moments. Restricted dynamic time warping is performed on this anchor point constraint cost matrix. During the recursive process, restricted dynamic time warping forcibly removes horizontal and vertical state transitions, requiring the time axis to monotonically advance to avoid matching to static images, thus eliminating the cumulative drift problem of traditional algorithms during the optimization phase. Through the anchor point constraints of the anchor point constraint cost field, the error of the highlight moments with the highest commercial value is ensured to be controlled within the frame-level range and is not affected by editing such as intros and outros, while retaining the flexible matching capability for non-highlight regions, avoiding matching failures when facing non-linear editing and visual secondary creation interference. Furthermore, semantic feature extraction is performed on the first frame sequence and the second frame sequence using a multi-stream semantic feature extraction method to enhance the recognition of the main content and resist background noise. We perform local slope analysis on the optimal path output by the constrained dynamic time warping, and deduce the playback strategy of the target video in reverse based on the local slope.

[0219] Based on the above embodiments Figure 1 The present application discloses a video processing method and a corresponding video processing apparatus, such as... Figure 4 As shown, the video processing device includes:

[0220] The acquisition unit 401 is used to acquire the first frame sequence of the source video and the second frame sequence of the target video;

[0221] Extraction unit 402 is used to extract semantic features from the first frame sequence to obtain the first semantic feature vector sequence of the source video, and to extract semantic features from the second frame sequence to obtain the second semantic feature vector sequence of the target video.

[0222] Construction unit 403 is used to construct a basic distance matrix using a first semantic feature vector sequence and a second semantic feature vector sequence;

[0223] The first determining unit 404 is used to determine the anchor point constraint cost matrix based on the base distance matrix and the predefined anchor point weight function;

[0224] The second determining unit 405 is used to determine the optimal path based on the anchor point constraint cost matrix.

[0225] The analysis and derivation unit 406 is used to analyze the optimal path to obtain the analysis results, and to deduce the playback strategy of the target video in reverse based on the analysis results.

[0226] Furthermore, the acquisition unit 401 includes:

[0227] The sampling module is used to perform sparse sampling on the acquired source video and target video;

[0228] The filtering module is used to execute a dynamic time warping algorithm to filter out all matching segments;

[0229] The operation module is used to perform operations on all matching segments to form a candidate time window for the target video;

[0230] The normalization module is used to perform frame rate recovery sampling within the candidate time window to obtain the source video after frame rate recovery and normalize it to obtain the first frame sequence, and to obtain the second frame sequence of the target video after frame rate recovery and normalize it to obtain the second frame sequence.

[0231] The first semantic feature vector sequence includes the first visual feature vector and the first spectral feature vector of the source video. The extraction unit 402 includes:

[0232] The first extraction module is used to extract the first global semantic vector and the first main semantic vector from the first frame sequence;

[0233] The first splicing processing module is used to splice the first global semantic vector and the first main semantic vector to generate the first visual feature vector sequence of the source video;

[0234] The second extraction module is used to extract audio features from the corresponding audio stream in the normalized first frame sequence to obtain the first spectral feature vector sequence of the source video.

[0235] Furthermore, the second semantic feature vector sequence includes the second visual feature vector and the second spectral feature vector of the target video. Extraction unit 402 includes:

[0236] The third extraction module is used to extract the second global semantic vector and the second main semantic vector from the second frame sequence;

[0237] The second splicing processing module is used to splice the second global semantic vector and the second main semantic vector to obtain the second visual feature vector sequence of the target video.

[0238] The fourth extraction module is used to extract audio features from the corresponding audio stream in the normalized second frame sequence to obtain the second spectral feature vector sequence of the target video.

[0239] Furthermore, building unit 403 includes:

[0240] The first calculation module is used to calculate the cosine distance between the first semantic feature vector sequence and the second semantic feature vector sequence;

[0241] The first overlay module is used to overlay the anchor point constraint cost matrix on the basis of calculating the base distance matrix;

[0242] The acquisition module is used to acquire the length of the first frame sequence of the source video and the length of the second frame sequence of the target video;

[0243] The second calculation module is used to calculate the basic distance matrix between the source video and the target video based on the cosine distance, the length of the first frame sequence, and the length of the second frame sequence.

[0244] Furthermore, the first determining unit 404 includes:

[0245] The second overlay module is used to overlay an anchor weight mask based on a flat-top Gaussian distribution on the basis of the calculated base distance matrix; wherein the anchor weight mask includes reward weights and penalty weights;

[0246] The application module is used to apply reward weights at preset positions corresponding to the target time in the source video and apply penalty weights in regions corresponding to non-target times in the source video to obtain the anchor point constraint cost matrix.

[0247] Furthermore, the second determining unit 405 includes:

[0248] The initialization module is used to initialize the cumulative cost of the first row of the anchor constraint cost matrix.

[0249] The third calculation module is used to perform state transition calculations on the initialized anchor point constraint cost matrix based on preset slope constraint conditions, and obtain the cumulative cost matrix.

[0250] The search module is used to find the minimum point in the last row of the cumulative cost matrix as the endpoint.

[0251] The backtracking module is used to backtrack the path backward from the endpoint in the cumulative cost matrix to obtain the optimal path.

[0252] Furthermore, the local slope includes a first slope and a second slope. The analysis and derivation unit 406 includes:

[0253] The determination module is used to determine the cutting point;

[0254] The segmentation module is used to segment the optimal path by cutting points to obtain the first and second path segments.

[0255] The fourth calculation module is used to calculate the first slope of the first path segment and the second slope of the second path segment.

[0256] The analysis module is used to perform segmented speed analysis based on the first slope, the second slope, and the preset slope value to obtain the speed ratio of the first path segment and the speed ratio of the second path segment, thereby completing the process of obtaining the playback strategy of the target video.

[0257] Furthermore, the video processing device includes:

[0258] The fifth calculation module is used to calculate the element standard deviation of the first semantic feature vector sequence and / or the second semantic feature vector sequence, respectively.

[0259] The comparison module is used to compare the standard deviation of elements with a preset threshold.

[0260] The judgment and marking module is used to determine and mark the corresponding image frame as an invalid frame with no information content if the standard deviation of an element is less than a preset threshold.

[0261] The sixth calculation module is used to adaptively adjust the visual and audio weights of invalid frames when calculating the basic distance matrix between the source video and the target video.

[0262] Furthermore, the video processing device includes:

[0263] The update module is used to adaptively update the penalty intensity and bandwidth parameters of the flat-top Gaussian distribution according to the closed-loop adaptive feedback mechanism.

[0264] The beneficial effects of this device embodiment are the same as those of the above method embodiment, and can be referred to accordingly. They will not be repeated here.

[0265] This application also provides a storage medium that includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to perform the video processing method described above.

[0266] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 5 As shown, it specifically includes a memory 501 and one or more instructions 502, wherein one or more instructions 502 are stored in the memory 501 and configured to be executed by one or more processors 503 to perform the above-mentioned video processing method.

[0267] The steps in the methods of the various embodiments of this application can be adjusted, combined, and deleted according to actual needs. It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0268] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0269] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A video processing method, characterized in that, The method includes: Obtain the first frame sequence of the source video and the second frame sequence of the target video; Semantic features are extracted from the first frame sequence to obtain a first semantic feature vector sequence of the source video, and semantic features are extracted from the second frame sequence to obtain a second semantic feature vector sequence of the target video. A basic distance matrix is ​​constructed using the first semantic feature vector sequence and the second semantic feature vector sequence; The anchor point constraint cost matrix is ​​determined based on the base distance matrix and the predefined anchor point weight function; The optimal path is determined based on the anchor point constraint cost matrix. The optimal path is analyzed to obtain the analysis results, and the playback strategy of the target video is deduced in reverse based on the analysis results.

2. The method according to claim 1, characterized in that, Obtain the first frame sequence of the source video and the second frame sequence of the target video, including: The acquired source and target videos are sparsely sampled. Execute the dynamic time warping algorithm to filter out all matching segments; Perform operations on all matching segments to form candidate time windows for the target video; Within the candidate time window, a frame rate recovery sampling operation is performed to obtain the source video after frame rate recovery and normalize it to obtain the first frame sequence. Then, the second frame sequence of the target video after frame rate recovery is obtained and normalized to obtain the second frame sequence.

3. The method according to claim 1, characterized in that, The first semantic feature vector sequence includes a first visual feature vector sequence and a first spectral feature vector sequence of the source video, and the second semantic feature vector sequence includes a second visual feature vector sequence and a second spectral feature vector sequence of the target video. Semantic feature extraction is performed on the first frame sequence to obtain the first semantic feature vector sequence of the source video, and semantic feature extraction is performed on the second frame sequence to obtain the second semantic feature vector sequence of the target video, including: Extract the first global semantic vector and the first main semantic vector from each frame of the first frame sequence; The first global semantic vector and the first main semantic vector are concatenated to generate the first visual feature vector sequence of the source video; Audio features are extracted from the audio stream corresponding to the first frame sequence to obtain the first spectral feature vector sequence of the source video; Extract the second global semantic vector and the second main semantic vector from the second frame sequence; The second global semantic vector and the second main semantic vector are concatenated to obtain the second visual feature vector sequence of the target video; Audio features are extracted from the corresponding audio stream in the second frame sequence to obtain the second spectral feature vector sequence of the target video.

4. The method according to claim 1, characterized in that, Constructing a basic distance matrix using the first semantic feature vector sequence and the second semantic feature vector sequence includes: Calculate the cosine distance between each feature vector in the first semantic feature vector sequence and each feature vector in the second semantic feature vector sequence; Obtain the length of the first frame sequence of the source video and the length of the second frame sequence of the target video; The fundamental distance matrix between the source video and the target video is calculated based on the cosine distance, the length of the first frame sequence, and the length of the second frame sequence.

5. The method according to claim 1, characterized in that, Based on the base distance matrix and the predefined anchor weight function, the anchor constraint cost matrix is ​​determined, including: An anchor weight mask based on a flat-top Gaussian distribution is superimposed on the calculated base distance matrix; wherein the anchor weight mask includes reward weights and penalty weights; The reward weight is applied at the preset position corresponding to the target time in the source video, and the penalty weight is applied in the region corresponding to the non-target time in the source video to obtain the anchor point constraint cost matrix.

6. The method according to claim 1, characterized in that, Determining the optimal path based on the anchor point constraint cost matrix includes: The first row of the anchor point constraint cost matrix is ​​initialized with cumulative cost. Based on the preset slope constraint conditions, the state transition calculation is performed on the initialized anchor point constraint cost matrix to obtain the cumulative cost matrix. Find the minimum point in the last row of the cumulative cost matrix as the endpoint; Starting from the endpoint, a reverse path backtracking is performed in the cumulative cost matrix to obtain the optimal path.

7. The method according to claim 1, characterized in that, The analysis results include a first slope and a second slope. The optimal path is analyzed to obtain the analysis results, and the playback strategy for the target video is derived in reverse based on these results, including: Determine the cutting point; The optimal path is divided by the cutting point to obtain the first path segment and the second path segment. Calculate the first slope of the first path segment, and calculate the second slope of the second path segment; Based on the first slope, the second slope, and the preset slope value, a segmented speed change analysis is performed to obtain the speed change ratio of the first path segment and the speed change ratio of the second path segment, thereby completing the process of obtaining the playback strategy of the target video.

8. The method according to claim 1, characterized in that, Also includes: Calculate the element standard deviations of the first semantic feature vector sequence and / or the second semantic feature vector sequence, respectively. The standard deviation of the element is compared with a preset threshold. If the standard deviation of the element is less than the preset threshold, the corresponding image frame is determined and marked as an invalid frame with no information content. When calculating the fundamental distance matrix between the source video and the target video, the visual and audio weights of invalid frames are adaptively adjusted.

9. The method according to claim 1, characterized in that, After analyzing the optimal path to obtain the analysis results, and then deriving the playback strategy of the target video based on the analysis results, the method further includes: Based on the closed-loop adaptive feedback mechanism, the penalty intensity and bandwidth parameters of the flat-top Gaussian distribution are adaptively updated.

10. A video processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire the first frame sequence of the source video and the second frame sequence of the target video. The extraction unit is used to extract semantic features from the first frame sequence to obtain the first semantic feature vector sequence of the source video, and to extract semantic features from the second frame sequence to obtain the second semantic feature vector sequence of the target video. The construction unit is used to construct a basic distance matrix using the first semantic feature vector sequence and the second semantic feature vector sequence; The first determining unit is used to determine the anchor point constraint cost matrix based on the base distance matrix and the predefined anchor point weight function. The second determining unit is used to determine the optimal path based on the anchor point constraint cost matrix. The analysis and derivation unit is used to analyze the optimal path to obtain the analysis result, and to deduce the playback strategy of the target video in reverse based on the analysis result.