Editing clip generation method and device based on video features, equipment and medium

Through video frame analysis, sparse tensor field and topological feature map processing, the problems of low efficiency and insufficient precision of traditional video editing are solved, and efficient and accurate video editing and content optimization are achieved.

CN120640062APending Publication Date: 2025-09-12PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510793545.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional video editing methods rely on manual labeling or fixed time window cutting, resulting in low editing efficiency and insufficient accuracy of the generated edited video clips. They are unable to cope with complex shot switching and scene changes, affecting the logic and viewing experience of the video.

Method used

By acquiring target video frames for analysis, generating sparse tensor fields, extracting global video features, constructing topological feature maps, performing cross-frame feature optimization and dynamic receptive field adjustment, optimizing preliminary clips, and finally generating a final clip set.

Benefits of technology

It improves the efficiency of video editing and the accuracy of generated video clips, enhances the coherence and logic of video content, accurately identifies key content areas, and avoids information omissions or redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640062A_ABST
    Figure CN120640062A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video feature-based edited fragment generation method, device, equipment and medium, and the method comprises the steps: obtaining a target to-be-edited video, carrying out the video frame analysis of the target to-be-edited video, obtaining a key video frame, and storing the key video frame in a database; generating a sparse tensor field according to the key video frame, extracting global video features of the target video to be edited by using the sparse tensor field, constructing a topological feature map according to the global video features, performing cross-frame feature optimization on the topological feature map to obtain a depth feature map, and performing dynamic receptive field adjustment on the depth feature map to obtain a target video to be edited. And obtaining a plurality of preliminary edited fragments, and performing content optimization on the preliminary edited fragments to obtain a final edited fragment set. According to the invention, the editing efficiency of the video and the precision of the generated edited video clip can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method, device, equipment and medium for generating clips based on video features. Background Art

[0002] Video editing technology plays a key role in a variety of fields, including film and television production, social media content generation, advertising and marketing, and education and training. The demand for automated editing is particularly urgent given the rapid rise of short video platforms. Traditional video editing methods, which mostly rely on manual annotation or fixed time window cutting, are not only inefficient and highly subjective, but also struggle to cope with frequent camera cuts and complex scene changes, resulting in videos that lack logic and visual appeal.

[0003] In the field of medical health, for example, it can be used to automatically process surgical videos, remote consultations, or rehabilitation training. However, because traditional editing methods rely on manual labeling or fixed-time cutting, it is difficult to accurately identify key medical operations and important doctor-patient interactions. Faced with problems such as frequent camera switching and complex environments in medical scenarios, the generated video content can easily lack professional logic and clinical reference value.

[0004] In the field of financial technology business, it can be applied to the compilation and publication of content such as financial live broadcasts, investment courses or user operation guides. However, traditional editing methods rely on manual and time-based processing, and cannot effectively adapt to the characteristics of information-intensive and fast-changing topics in financial videos. As a result, the editing results are logically confusing and the focus is unclear, affecting the professionalism and dissemination effect of the content.

[0005] In summary, traditional editing technology mainly relies on manual labeling or cutting based on fixed time windows. The operation is cumbersome and the processing efficiency is low. In addition, it lacks accuracy when identifying key content and dealing with complex shot switching, which often leads to inaccurate selection of video clips, affecting the coherence and expression effect of the content.

[0006] Therefore, the current technology has the problems of low efficiency in video editing and low accuracy in the generated edited video segments. Summary of the Invention

[0007] The present invention provides a method, device, equipment and medium for generating clip segments based on video features, the main purpose of which is to solve the problems of low video clipping efficiency and low precision of generated clipped video segments.

[0008] In a first aspect, to achieve the above-mentioned purpose, the present invention provides a method for generating clips based on video features, comprising:

[0009] Obtaining a target video to be edited, performing video frame analysis on the target video to be edited, and obtaining key video frames;

[0010] Generating a sparse tensor field according to the key video frame, and extracting global video features of the target video to be edited using the sparse tensor field;

[0011] Constructing a topological feature map according to the global video features;

[0012] Performing cross-frame feature optimization on the topological feature map to obtain a depth feature map;

[0013] Dynamically adjusting the receptive field of the depth feature map to obtain a number of preliminary clips;

[0014] Content optimization is performed on the preliminary clips to obtain a final clip set.

[0015] In a second aspect, the present invention further provides a video feature-based clip generation device, comprising:

[0016] A video key frame extraction module is used to obtain a target video to be edited, perform video frame analysis on the target video to be edited, and obtain key video frames;

[0017] A sparse tensor field construction module, configured to generate a sparse tensor field according to the key video frames, and extract global video features of the target video to be edited using the sparse tensor field;

[0018] A topological feature map construction module, configured to construct a topological feature map according to the global video features;

[0019] A cross-frame feature optimization module is used to perform cross-frame feature optimization on the topological feature map to obtain a depth feature map;

[0020] A dynamic receptive field adjustment module, configured to dynamically adjust the receptive field of the depth feature map to obtain a plurality of preliminary clips;

[0021] The clip optimization module is used to optimize the content of the preliminary clips to obtain a final clip set.

[0022] In a third aspect, the present invention further provides an electronic device, comprising:

[0023] at least one processor; and,

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the above-mentioned method for generating clips based on video features.

[0026] In a fourth aspect, the present invention further provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned method for generating clips based on video features.

[0027] The present invention obtains a target video to be edited, performs video frame analysis on the target video to be edited, obtains key video frames, automatically classifies the video into a text video or an image video according to a text quantity threshold, adopts differentiated key frame extraction strategies for different types, effectively improves the accuracy and representativeness of the key video frames, generates a sparse tensor field based on the key video frames, and uses the sparse tensor field to extract the global video features of the target video to be edited, accurately screens out sparse coordinate points representing significant edge and texture features, effectively captures the local details and structural information of the video, and can generate global features reflecting the overall content and semantics of the video through aggregation of the sparse tensor field, thereby improving the accuracy of video analysis, and constructs a topological feature map based on the global video features, effectively capturing the spatial structure and temporal sequence of the video content. Relationship, while fusing global video features, to achieve a unified expression of local details and overall semantics of the video, perform cross-frame feature optimization on the topological feature map to obtain a depth feature map, and achieve deep mining and optimization of spatiotemporal correlations between nodes in the topological feature map, effectively enhancing the fusion and expression capabilities of cross-frame features, and dynamically adjust the receptive field of the depth feature map to obtain several preliminary clips, which can accurately identify semantically rich and compact content areas in the video, effectively avoiding information omissions or redundancy caused by a fixed range, and optimize the content of the preliminary clips to obtain a final clip set, and merge adjacent clips based on time intervals and semantic similarity to improve the continuity and overall coherence of the clips, which can effectively improve the efficiency of video editing and the accuracy of generated clipped video clips. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0029] Figure 1A schematic diagram of an application environment of a method for generating clips based on video features according to an embodiment of the present invention;

[0030] Figure 2 A schematic flow chart of a method for generating clips based on video features provided by one embodiment of the present invention;

[0031] Figure 3 A schematic diagram of a flow chart of a sparse tensor field construction module in a method for generating video clips based on video features provided by one embodiment of the present invention;

[0032] Figure 4 A schematic diagram of modules of a video feature-based clip generation device provided by one embodiment of the present invention;

[0033] Figure 5 A schematic structural diagram of an electronic device for implementing a method for generating clips based on video features provided by an embodiment of the present invention;

[0034] Figure 6 Another structural diagram of an electronic device for implementing a method for generating clips based on video features provided by an embodiment of the present invention.

[0035] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] The embodiment of the present application provides a method for generating clips based on video features, and the execution subject of the method for generating clips based on video features includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the device provided by the embodiment of the present application. In other words, the method for generating clips based on video features can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0039] The embodiment of the present invention provides a method for generating clips based on video features, which can be applied in the following situations: Figure 1In the application environment. Among them, the client communicates with the server through the network. The server can obtain the target video to be edited through the client, perform video frame analysis on the target video to be edited, obtain key video frames, and automatically classify the video into text video or image video according to the text quantity threshold. Differentiated key frame extraction strategies are adopted for different types, which effectively improves the accuracy and representativeness of key video frames, generates a sparse tensor field based on the key video frames, and uses the sparse tensor field to extract the global video features of the target video to be edited, accurately screens out sparse coordinate points representing significant edge and texture features, and effectively captures the local details and structural information of the video. By aggregating the sparse tensor field, it can generate global features that reflect the overall content and semantics of the video, thereby improving the accuracy of video analysis. A topological feature map is constructed based on the global video features to effectively capture the spatial structure and temporal relationship of the video content, while integrating the entire The local video features are used to achieve a unified expression of local details and overall semantics of the video. The topological feature map is optimized across frames to obtain a deep feature map, which achieves deep mining and optimization of the spatiotemporal correlation between nodes in the topological feature map, effectively enhancing the fusion and expression capabilities of cross-frame features. The deep feature map is dynamically adjusted for receptive field to obtain several preliminary clips, which can accurately identify semantically rich and compact content areas in the video, effectively avoiding information omission or redundancy caused by a fixed range. The preliminary clips are optimized for content to obtain a final clip set. Adjacent clips are merged based on time interval and semantic similarity to improve the continuity and overall coherence of the clips, which can effectively improve the efficiency of video editing and the accuracy of the generated clipped video clips. Finally, the final clip set is output and fed back to the user client. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0040] The following is an explanation of the specification of the present invention. The present invention uses sparse tensor fields to accurately screen out sparse coordinate points representing significant edge and texture features, effectively capturing local details and structural information of the video. By aggregating the sparse tensor fields, it can generate global features that reflect the overall content and semantics of the video, thereby improving the accuracy of video analysis. The topological feature map effectively captures the spatial structure and temporal relationship of the video content, while integrating global video features to achieve a unified expression of local details and overall semantics of the video, thereby improving the efficiency and accuracy of video editing.

[0041] Reference Figure 2FIG. 1 is a flow chart of a method for generating a clip based on video features according to an embodiment of the present invention. In this embodiment, the method for generating a clip based on video features includes:

[0042] S1. Obtain a target video to be edited, perform video frame analysis on the target video to be edited, and obtain key video frames.

[0043] In this embodiment of the present invention, a target video is acquired and analyzed frame by frame, with OCR used to determine the video type. For videos with text, the number of characters in each frame is calculated after extraction, and frames with a higher-than-average character count are selected as key frames. For videos with only images, inter-frame image changes are determined based on Structural Image Similarity (SSIM), and the first frame with a significant decrease in similarity is recorded as the key frame, thereby accurately extracting key content.

[0044] In specific healthcare scenarios, in medical imaging explanation videos or remote surgery training videos, the system is used to automatically filter out images containing important medical explanatory text (such as diagnostic recommendations, step-by-step instructions, etc.) or key surgical operation changes. For example, in medical videos with explanation subtitles or graphic annotations, the system can extract key explanation content by identifying text-intensive frames; in text-free surgical records, the system can detect key operation steps through image changes, providing medical teaching systems with concise images, improving learning efficiency and content retrieval capabilities.

[0045] In specific FinTech scenarios, in videos demonstrating financial products, interpreting data, or explaining reports, the system assists in identifying key frames that display core indicators, risk warnings, or concluding statements. For analytical videos with charts and explanations, text recognition is used to extract frames with high information density, enabling concise report summaries. In videos demonstrating transaction processes or with frequent interface changes, key operational steps are detected through screen changes, helping to generate transaction reviews, automated risk management records, or audit clips, enhancing intelligent analysis and regulatory support capabilities for financial content.

[0046] In an embodiment of the present invention, the step of performing video frame analysis on the target video to be edited to obtain key video frames includes:

[0047] Decoding the target video to be edited to obtain a plurality of initial video frames;

[0048] Performing optical character recognition on the characters in the initial video frame and counting the characters to obtain a total number of the characters;

[0049] Determine whether the total number of characters is greater than a preset threshold;

[0050] If the total number of characters is greater than the number threshold, the target video to be edited is determined to be a text video type;

[0051] Performing valid text recognition on the initial video frame according to the text video type to obtain the number of valid texts in each frame;

[0052] Calculating an average value of all the valid character numbers, and screening out initial video frames corresponding to valid character numbers greater than the average value;

[0053] Using the filtered initial video frame as a key video frame;

[0054] If the total number of characters is less than or equal to the number threshold, the target video to be edited is determined to be an image video type;

[0055] Randomly select two initial video frames with consecutive preset first timestamps as a group of frames to be detected;

[0056] Calculating the image similarity of the frames to be detected according to the image and video type, and screening out a first group of target detection frames whose image similarity is greater than a preset similarity threshold in a preset order;

[0057] An initial video frame with a larger preset first timestamp among the filtered target detection frames is used as a key video frame.

[0058] In detail, the video file is parsed, the compressed and encoded data is restored to the original image data, and then each frame of the video is extracted in turn to obtain several initial video frames arranged in chronological order, providing basic materials for subsequent content analysis and editing processing.

[0059] The image is preprocessed, such as grayscale conversion, denoising, and edge enhancement, to improve recognition accuracy. The optical character recognition (OCR) algorithm is used to scan the image area, that is, the sliding window or image segmentation algorithm is used to locate the area that may contain text. Feature extraction is performed on these candidate areas, and the graphic features such as lines, contours, and textures are analyzed. These image features are converted into characters through deep learning models or template matching methods. The recognized text content is extracted frame by frame and statistically summarized to finally obtain the total number of words in the video.

[0060] After completing the recognition and counting of the text in the video frame, the total number of characters obtained is compared with the quantity threshold. If the total number of recognized characters is greater than the quantity threshold, the target video to be edited is determined to be a text video type, that is, a video with text information as the main content feature.

[0061] The initial video frames are analyzed frame by frame, using optical character recognition (OCR) technology to accurately identify text areas within each frame. Each frame undergoes text area detection, followed by character recognition and filtering within the detected areas to remove meaningless characters, repetitive content, or noise interference. Finally, the number of valid text characters in each frame is counted. The number of valid text characters across all frames is aggregated and averaged to obtain a text density benchmark for the video frame sequence. The number of valid text characters within each frame is then compared and judged. Frames with a greater than average value are selected and marked as key video frames, deemed more representative in terms of information density and valuable for editing.

[0062] If the total number of characters is less than or equal to the quantity threshold, the target video to be edited is judged as an image video type, and two adjacent initial video frames with continuous timestamps are randomly selected on the video timeline. These two frames are combined into a group of frames to be detected for subsequent image content analysis and change detection.

[0063] For image and video types, the similarity between two frames of each group of frames to be detected is calculated. The color, texture and structural features of the images are compared to quantify the similarity values ​​by using methods such as cosine similarity, structural similarity index SSIM or mean square error MSE. The first group of target detection frames with a similarity greater than the similarity threshold are screened out. It is considered that this group of frames has small changes, and the initial video frame with a larger timestamp in this group is further selected as the key video frame.

[0064] Videos are automatically categorized as text or image based on a threshold for the number of words. Differentiated keyframe extraction strategies are employed for each type, effectively improving the accuracy and representativeness of key video frames. For text videos, precise text recognition and statistical analysis are used to filter out information-dense frames, ensuring the logical coherence of the content. For image videos, image similarity analysis is used to filter out frames with significant changes, avoiding redundant content and improving editing efficiency and video quality. This achieves intelligent discrimination and refined processing within the automated editing process.

[0065] S2. Generate a sparse tensor field according to the key video frame, and use the sparse tensor field to extract global video features of the target video to be edited.

[0066] In this embodiment of the present invention, the gradient magnitude maps of the selected keyframes are calculated and sparse coordinate points with significant changes are extracted based on the 95th percentile. A lightweight feature extractor, combining MobileNetV3 and Swin Transformer, is then used to generate low-dimensional deep semantic features at these sparse locations, thereby constructing a sparse tensor field. The sparse tensor fields of multiple keyframes collectively constitute the global video features of the video, significantly reducing the computational effort while preserving the essential information content, providing efficient support for video editing, understanding, and analysis tasks.

[0067] In specific healthcare scenarios, this technology is used for surgical process recording, remote consultations, or medical teaching videos. Through keyframe screening and sparse tensor construction, it can accurately extract key visual features such as organ exposure, surgical incisions, and operation steps, avoiding frame-by-frame processing of the entire video. This significantly reduces computational complexity while retaining core medical information, facilitating the generation of surgical summaries, training AI diagnosis and treatment models, or building standardized teaching materials.

[0068] In specific FinTech scenarios, this approach is used in market analysis, trading process demonstrations, or financial education videos. By extracting keyframes containing chart changes, data analysis, and risk warnings, and constructing sparse tensor fields for regions of significant change, it can efficiently extract deep financial semantic features, enabling automated video summarization, intelligent tag recommendations, and screening for key compliance content.

[0069] Figure 3 A schematic flow chart of a sparse tensor field construction module in a method for generating clips based on video features provided by one embodiment of the present invention.

[0070] In an embodiment of the present invention, generating a sparse tensor field according to the key video frames and extracting global video features of the target video to be edited using the sparse tensor field includes:

[0071] Obtaining a key video image corresponding to the key video frame, and calculating a gradient magnitude map of the key video image;

[0072] Screening out gradient amplitudes greater than a preset amplitude threshold in the gradient amplitude map, and obtaining a sparse coordinate set corresponding to the screened gradient amplitudes;

[0073] Performing hybrid feature extraction on the sparse coordinate set to obtain a depth feature vector;

[0074] constructing a sparse tensor field according to the depth feature vector;

[0075] The sparse tensor field is aggregated to obtain global video features.

[0076] Specifically, the corresponding key video image is extracted from the key video frame, and then the key video image is processed by a gradient calculation algorithm (such as the Sobel operator or the Scharr operator). The gradient values ​​of each pixel in the key video image in the horizontal and vertical directions are calculated. The gradient information in these two directions is combined to generate a gradient amplitude map reflecting the intensity of changes in image edges and details.

[0077] Compare each pixel value in the gradient magnitude map point by point, and filter out all gradient magnitudes greater than the magnitude threshold. These larger gradient magnitudes usually correspond to significant edges or texture details in the image. Record the pixel positions where these qualified gradient magnitudes are located to form a sparse coordinate set to represent the key point positions in the edge feature set in the image. where v t Indicates the t-th frame, T represents the total number of video frames, and N key video frames are filtered out (usually N<<T), the gradient amplitude map calculation formula is as follows:

[0078]

[0079] Among them, k i (x,y) represents the pixel coordinates of the key video image corresponding to the i-th key video frame, represents the gradient operator, G i (x,y) represents the gradient magnitude map, τ represents the magnitude threshold, Represents a sparse coordinate set.

[0080] Multidimensional mixed features are extracted from each point in the sparse coordinate set, including position coordinates, gradient amplitude, and local texture information of surrounding pixels. These features are fused and encoded through a deep learning model to generate deep feature vectors with rich semantic information. These deep feature vectors are mapped to three-dimensional or multi-dimensional space according to their spatial coordinates to construct a sparse tensor field, which is used to express the sparse but important structural and texture features in key video images. The calculation formula is as follows:

[0081]

[0082] Where Ψ represents a hybrid feature extraction extractor, f represents a deep feature vector, (x, y) represents the pixel coordinates of the key video image, and k i represents the key video image corresponding to the i-th key video frame, represents a sparse coordinate set, Represents a sparse tensor field.

[0083] By aggregating the deep feature vectors in the sparse tensor field, the scattered local feature information is integrated and weightedly fused using methods such as pooling or attention mechanism to extract the global feature representation representing the entire video content, thereby obtaining the global video features that reflect the overall structure and semantic information of the video.

[0084] By extracting image gradient information from key video frames, the team accurately screened sparse coordinate points representing significant edge and texture features. Deep encoding was then combined with multi-dimensional hybrid features to construct a sparse tensor field, effectively capturing the video's local details and structural information. By aggregating the sparse tensor field, global features reflecting the overall content and semantics of the video were generated, improving the accuracy and expressiveness of video analysis and providing a solid data foundation for subsequent intelligent editing and content understanding.

[0085] S3. Construct a topological feature map based on the global video features.

[0086] In an embodiment of the present invention, sparse feature points corresponding to global video features in all key frames are used as nodes to construct a node set. Within the same frame, spatial adjacency edges are established based on the spatial distance between feature points being less than a preset threshold. Between adjacent key frames, temporal association edges are established for feature points with similar spatial positions, thereby forming a topological feature graph that simultaneously reflects spatial locality and temporal continuity, providing a structural basis for subsequent feature propagation and video understanding.

[0087] In specific medical and health scenarios, by constructing the significant feature points of key frames in surgical or medical treatment videos into a topological map connected in space and time, it helps to capture continuous key actions and anatomical structure changes in surgical steps, achieve dynamic understanding and anomaly detection of complex medical operation processes, and assist doctors in accurate review and intelligent training.

[0088] In specific FinTech scenarios, by connecting key feature points that are adjacent in time and space in trading demonstrations or market interpretation videos, a topological structure that reflects operational processes and data fluctuations is established, supporting continuous tracking and pattern recognition of trading behaviors, and improving the accuracy and timeliness of risk monitoring, compliance review, and intelligent trading analysis.

[0089] In an embodiment of the present invention, constructing a topological feature map according to the global video features includes:

[0090] Using the sparse coordinate set of each key video frame as a graph node;

[0091] Randomly select two sparse coordinates in the sparse coordinate set as coordinate pairs to be traversed;

[0092] Calculating the two-dimensional space distance of each of the coordinate pairs to be traversed;

[0093] Determining whether the two-dimensional space distance is less than a preset plane distance threshold;

[0094] If the two-dimensional space distance is greater than or equal to the plane distance threshold, no undirected graph edge is established in the coordinate pair to be traversed;

[0095] If the two-dimensional space distance is less than the plane distance threshold, establishing an undirected graph edge in the coordinate pair to be traversed;

[0096] Randomly select two key video frames with consecutive preset second timestamps as adjacent key frames;

[0097] Calculating the three-dimensional space distance between every two sparse coordinates in the adjacent key frames;

[0098] Determining whether the three-dimensional space distance is less than a preset timing threshold;

[0099] If the three-dimensional spatial distance is greater than or equal to the timing threshold, no cross-frame connection edge is established in the adjacent key frames;

[0100] If the three-dimensional spatial distance is less than the temporal threshold, establishing a cross-frame connection edge in the adjacent key frames;

[0101] A topological feature graph is constructed according to the graph nodes, the undirected graph edges, the cross-frame connection edges and the global video features.

[0102] In detail, two coordinate points are randomly selected from the sparse coordinate set to form a coordinate pair to be traversed. The Euclidean distance between the two points on the two-dimensional plane is calculated to quantify the spatial interval between them. The system compares the calculated distance with the plane distance threshold. If the distance is greater than or equal to the plane distance threshold, it indicates that the two points are far apart and do not have direct spatial proximity. Therefore, no undirected graph edge will be established between the two coordinate points; if the distance is less than the plane distance threshold, it means that the two points are relatively close in space. The system will establish an undirected graph edge between the pair of coordinate points to reflect the spatial connection between them and construct a local topological structure between the nodes.

[0103] Two key video frames with adjacent timestamps are randomly selected as adjacent key frames. For each pair of sparse coordinate points in these two frames, the distance between the adjacent key frames in three-dimensional space is calculated. The distance comprehensively considers the differences in two-dimensional spatial coordinates and time dimension. The calculated three-dimensional distance is compared with the timing threshold. If the distance is greater than or equal to the timing threshold, it indicates that the two points are far apart in space or time, and no cross-frame connection edge is established; if the distance is less than the timing threshold, the system establishes a cross-frame undirected connection edge between the corresponding sparse coordinate points of the adjacent key frames, reflecting the spatial proximity and temporal continuity between the two frames.

[0104] Based on graph nodes (i.e., sparse coordinate points in key video frames), undirected graph edges (representing connections between spatially adjacent nodes in the same frame), cross-frame connection edges (reflecting the spatiotemporal association between adjacent key frames), and global video features, these elements are organically integrated to construct a topological feature graph containing spatial and temporal information. It not only reflects the local and cross-frame relationships between nodes, but also combines global video semantic information, providing a rich and multi-dimensional expression basis for subsequent graph neural network analysis and video content understanding.

[0105] By combining the spatial position and temporal sequence of sparse coordinate points in key video frames, the team meticulously constructs undirected graph edges within the same frame and connecting edges across frames, effectively capturing the spatial structure and temporal relationships of video content. This, combined with global video features, enables a unified representation of both local details and overall video semantics. Topological feature graphs not only enhance the expressive power and structural integrity of video features but also provide rich, multidimensional information support for subsequent graph-based deep learning models, thereby enhancing the efficiency and precision of video analysis, understanding, and intelligent editing.

[0106] S4. Perform cross-frame feature optimization on the topological feature map to obtain a depth feature map.

[0107] In an embodiment of the present invention, by introducing an improved graph attention network, multi-layer feature propagation and optimization are performed on the constructed topological feature graph: in each layer, the node calculates the attention weight of the neighboring nodes through a gating mechanism, aggregates the neighboring features by weight, and performs a nonlinear transformation after splicing them with its own features, and updates the feature representation layer by layer. After 2-3 layers of propagation, the node features fully integrate the spatial adjacency and cross-frame temporal context information, thereby forming a deeper feature graph with stronger expressive ability, providing a high-quality feature foundation for subsequent video understanding and processing.

[0108] In specific medical and health scenarios, it is used in surgical videos, endoscopic images or dynamic tracking of lesions. The graph attention mechanism is used to optimize the propagation of sparse features across frames in key areas, effectively integrating the contextual information of anatomical structure changes and continuous operation behaviors, thereby enhancing the accuracy and robustness of disease identification, surgical procedure analysis or remote assisted diagnosis.

[0109] In specific FinTech scenarios, it is applied to financial transaction demonstrations, risk control monitoring videos and other content. By optimizing the propagation of graph neural networks across frame feature points such as chart fluctuations and user operations, it can achieve accurate modeling of time-sensitive key behaviors, help automatically identify potential risk behaviors, build user operation chains or generate structured financial event labels, and enhance the capabilities of intelligent analysis and compliance review.

[0110] In an embodiment of the present invention, performing cross-frame feature optimization on the topological feature map to obtain a depth feature map includes:

[0111] Obtaining an initial attention weight, and using the initial attention weight to perform feature propagation and update on the topological feature map to obtain an updated feature map;

[0112] Calculating an updated attention weight between each pair of adjacent nodes in the updated feature graph;

[0113] The updated attention weight is used as a new initial attention weight, and the initial attention weight is returned to obtain the initial attention weight, and the topological feature map is propagated and updated using the initial attention weight to obtain the step of updating the feature map, and the number of returns is counted;

[0114] When the number of returns is equal to a preset number threshold, the return is ended, and the updated feature map returned for the last time is used as the depth feature map.

[0115] In detail, the attention weight is initialized for each node in the topological feature graph, which is usually obtained by calculating the similarity between the initial feature vector of the node and the features of the neighboring nodes. The graph attention network (GAT) mechanism is used to calculate the attention coefficient of each node and its neighboring nodes. This process includes linear transformation of node features, calculation of compatibility scores between nodes (such as dot products or learnable functions), and normalization through the softmax function to obtain normalized attention weights. According to these weights, the features of the neighboring nodes are weighted summed to achieve feature aggregation and update. For node u in the topological feature graph, the feature update formula at layer l is:

[0116]

[0117] Where a represents the first learnable parameter, W represents the second learnable parameter, represents the features of node u in the lth layer, represents the characteristics of node v, represents the feature of the adjacent node w of the l-th layer node u, T represents the transposition operation, LeakyReLU represents the activation function, σ represents the activation function, W (l) represents the second learnable parameter of the lth layer, CONCAT represents feature concatenation, represents the set of adjacent nodes of node u, represents the feature of node u in the lth layer, α uv represents the attention weight, Represents the characteristics of the adjacent node v of node u, Represents the features of node u in the l+1th layer. Through multi-layer stacking and multiple iterations, node features are continuously propagated and fused, and finally an updated feature graph containing context information and graph structure relationships is obtained.

[0118] Their updated feature vectors are linearly transformed respectively, and then the compatibility score of the features of the two is calculated (such as dot product or based on a learnable feedforward neural network). The scores of all adjacent nodes are then normalized using the softmax function to obtain the updated attention weights between each pair of adjacent nodes, reflecting the importance of information transfer between nodes.

[0119] The updated attention weight obtained from each calculation is fed back as the initial attention weight for the next round of feature propagation. The weight is used to propagate and update the topological feature map, generate a new updated feature map, and record the number of returns. The process is iterative. When the number of returns reaches the preset threshold, the iteration is stopped. Finally, the feature map obtained from the last update is output as a deep feature map to represent the multi-level semantic information and graph structure features of the node.

[0120] Through multiple rounds of feature propagation and updates based on the attention mechanism, we achieve deep mining and optimization of the spatiotemporal correlations between nodes in the topological feature graph, effectively enhancing the fusion and expression capabilities of cross-frame features. Continuously iteratively calculating and updating attention weights enables the model to dynamically adjust the importance of information transfer between nodes, capturing fine-grained changes and temporal dependencies in video content. The resulting deep feature graphs are enriched with semantic information and structural consistency, helping to improve the accuracy and effectiveness of applications such as video understanding, retrieval, and intelligent editing.

[0121] S5. Dynamically adjust the receptive field of the depth feature map to obtain several preliminary clips.

[0122] In an embodiment of the present invention, the DBSCAN clustering algorithm is used to cluster the node features in the depth feature map, extract the significant areas with semantic consistency, and calculate the semantic entropy in the area based on each cluster center. The receptive field radius is dynamically adjusted by the entropy value to achieve an adaptive response to the semantic complexity. The preliminary editing interval is determined according to the frame numbers corresponding to the nodes in each cluster, and the boundaries of the start and end frames are optimized by judging the completeness of the subtitles and the switching of the screen shots, and finally a clip with a reasonable structure and semantic coherence is generated.

[0123] In specific healthcare scenarios, the system is used for automatic editing and summarization of surgical records, endoscopic examinations, or lesion observation videos. Clustering identifies semantically consistent key operation areas and dynamically adjusts the receptive field based on the complexity of the operation, effectively extracting key diagnostic and treatment segments. Furthermore, through subtitle and shot boundary optimization, it ensures that the edited content is both medically critical and visually coherent, improving the quality of teaching materials, remote diagnosis, and AI-assisted report generation.

[0124] In specific FinTech scenarios, it is suitable for intelligent editing of financial live broadcasts, business process demonstrations, or risk control monitoring videos. It clusters to identify explanations or operation segments with strong continuity and semantic consistency, and dynamically adjusts the receptive field based on the density of charts or market fluctuation information, thereby extracting key content segments. The boundary optimization mechanism further ensures the integrity of subtitles and picture logic, providing efficient video clips for content summarization, compliance evidence collection, and knowledge base construction.

[0125] In an embodiment of the present invention, the dynamic receptive field adjustment of the depth feature map to obtain a plurality of preliminary clips includes:

[0126] Performing cluster analysis on the depth feature map to obtain several target clusters;

[0127] Determine the receptive field area according to the cluster center of the target cluster, and calculate the semantic entropy value of the receptive field area;

[0128] Dynamically adjusting the radius of the receptive field area according to the semantic entropy value to obtain an updated cluster;

[0129] Obtaining frame numbers corresponding to the graph nodes in the update cluster in the target video to be edited;

[0130] The smallest frame number is used as the clip start frame, and the largest frame number is used as the clip end frame;

[0131] Obtaining a search frame distance, adjusting boundaries of adjacent frames of the clip start frame according to the search frame distance, and updating the start frame;

[0132] A plurality of initial clipping segments are generated according to the update start frame and the clipping end frame.

[0133] In detail, cluster analysis is performed on the node features in the deep feature graph, usually using K-means, spectral clustering or density-based clustering algorithms, to automatically group nodes with similar feature representations into the same cluster, and then divide several target clusters. These target clusters reflect the content areas in the video that are highly correlated in terms of time, space and semantics.

[0134] Taking the cluster center of each target cluster as the base point, combined with the node features within a certain range around the cluster center, the corresponding receptive field area is determined. The receptive field area covers the representative spatiotemporal features within the cluster. The semantic distribution of the node features within the receptive field area is calculated, and its semantic entropy value is evaluated based on the information entropy principle to quantify the complexity and diversity of the semantic information in the area, and assist in judging the importance and distinctiveness of the area in the video content.

[0135] Based on the semantic entropy calculated for the receptive field, the radius of that region is dynamically adjusted. When the semantic entropy is high, indicating rich and complex information within the region, the receptive field radius is appropriately expanded to cover more relevant features. When the semantic entropy is low, indicating a region with relatively simple or redundant semantics, the radius is reduced to focus on the core information. This dynamic adjustment process updates the scope of the target cluster to more accurately reflect the semantic structure and spatiotemporal distribution of the video content.

[0136] Extract the target video frame numbers to be edited corresponding to all graph nodes from the update cluster, traverse and determine the minimum and maximum values ​​of these frame numbers, and use them as the start and end frames of the video clip, respectively, to ensure that the clip completely covers the key content interval represented by the cluster.

[0137] Get the preset search frame distance parameters, use them as a benchmark to adjust the boundaries of the adjacent frames of the clip start frame, extend the start frame position forward or backward to ensure that the key content is completely included, and use the adjusted start frame and the original clip end frame as the interval boundaries to divide the interval into several initial clip segments, providing basic materials for subsequent fine editing and content optimization.

[0138] By performing cluster analysis on deep feature maps and adjusting the dynamic receptive field based on semantic entropy, it is possible to accurately identify semantically rich and compact content areas in the video, effectively avoiding information omission or redundancy caused by a fixed range; by adjusting the boundaries of the starting frames of the clips, the integrity and continuity of the clips are ensured. The final generated preliminary clips are more in line with the spatiotemporal distribution and semantic characteristics of the video content, which helps to improve the accuracy of automatic editing and the viewing experience.

[0139] S6. Optimize the content of the preliminary clips to obtain a final clip set.

[0140] In an embodiment of the present invention, low-quality or falsely detected segments are eliminated, and their validity is judged based on indicators such as segment length, semantic entropy, degree of picture change, and text information. Semantic similarity analysis and temporal continuity judgment are performed on adjacent segments, and semantically related clip segments are automatically merged to improve the integrity and coherence of the video content, ultimately generating a final set of clips with a reasonable structure and clear semantics.

[0141] In specific medical and health scenarios, it is used for editing and refining surgical records, clinical operation teaching, or lesion observation videos, automatically eliminating invalid or repeated segments, such as preoperative preparation, static images, and other low-information segments, and merging continuous operation links through semantic relevance, making the final edited segments more medically valuable and teaching coherent.

[0142] In specific FinTech scenarios, content optimization for live financial teaching, risk control demonstrations, or business process recordings automatically filters out invalid speech fragments or pauses, retaining core business operations and key instructions; at the same time, semantically related explanation fragments are merged to form a more complete and logically clear financial event chain or knowledge module.

[0143] In an embodiment of the present invention, optimizing the content of the preliminary clips to obtain a final clip set includes:

[0144] Calculating the average semantic entropy of the video frames in the preliminary clip one by one;

[0145] Determining whether the average semantic entropy is greater than a preset average entropy value threshold;

[0146] If the average semantic entropy is less than or equal to the average entropy threshold, deleting the preliminary clipping segments corresponding to the average semantic entropy less than or equal to the average entropy threshold;

[0147] If the average semantic entropy is greater than the average entropy value threshold, the preliminary clipping segment corresponding to the average semantic entropy greater than the average entropy value threshold is used as the clipping retention segment;

[0148] Randomly selecting two clips of the clip with adjacent preset third timestamps as a clip comparison group;

[0149] calculating the time interval and semantic similarity between the comparison groups of the fragments;

[0150] If the time interval is less than a preset time threshold and the semantic similarity is greater than a preset semantic similarity threshold, merging the segment comparison groups to obtain an updated clip;

[0151] The retained clips and the updated clips are aggregated to obtain a final clip set.

[0152] In detail, the video frames in each preliminary clip are traversed one by one to obtain the semantic entropy value of the feature area corresponding to each frame. Then, the average semantic entropy of all frames in the preliminary clip is calculated to quantify the overall semantic complexity and information richness of the clip, providing a basis for subsequent clip optimization and content screening.

[0153] The average semantic entropy of each preliminary clip is compared with the average entropy threshold. If the average semantic entropy of a preliminary clip is less than or equal to the average entropy threshold, it is considered that the semantic information of the clip is relatively monotonous or lacks sufficient content complexity. Therefore, the clip is deleted from the preliminary clip list to avoid invalid or redundant content affecting the editing quality; if the average semantic entropy of a clip is greater than the average entropy threshold, it is considered that the clip is semantically rich and has high information value. It is then marked and added to the clip retention set to ensure that the subsequent processing stage focuses on retaining these video clips with high information density and viewing value, thereby improving the content quality of the final editing effect.

[0154] For each pair of adjacent preliminary clips in the retained clip set, the time interval between them is calculated—that is, the difference between the start frame number of the latter clip and the end frame number of the previous clip. Based on a deep feature vector or semantic representation, the semantic similarity between the two clips is calculated to quantify the degree of content relevance. If the time interval between two adjacent preliminary clips is less than a time threshold and the semantic similarity is greater than a similarity threshold, indicating that the two clips have strong continuity and relevance in both time and content, the system then merges the two clips to generate a new, updated clip. This merging process includes combining the frame intervals and feature information of the two clips. Finally, the updated clip replaces the original two clips in the retained clip set, completing the update of the set. This results in a final clip set that is more compact, coherent, non-redundant, and easy to watch.

[0155] By calculating the average semantic entropy of the preliminary clips, we can effectively screen out clips with low information content and high redundancy, ensure the quality of the clip content and the richness of expression, merge adjacent clips based on time interval and semantic similarity, improve the continuity and overall coherence of the clips, and avoid unnecessary segmentation. Finally, through the combination of content optimization and merging steps, the generated clip set is both refined and complete, greatly improving the efficiency and accuracy of video editing.

[0156] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0157] like Figure 4 FIG. 1 is a functional module diagram of a video feature-based clip generation device provided by an embodiment of the present invention.

[0158] In the embodiment of the present disclosure, a video feature-based clip generation device is provided, and the video feature-based clip generation device corresponds to the video feature-based clip generation method of the above embodiment. Figure 4As shown, the video feature-based clip generation device 100 can be installed in an electronic device. According to the functions to be implemented, the video feature-based clip generation device 100 includes a video key frame extraction module 101, a sparse tensor field construction module 102, a topological feature map construction module 103, a cross-frame feature optimization module 104, a dynamic receptive field adjustment module 105, and a clip optimization module 106. The functional modules are described in detail as follows:

[0159] The video key frame extraction module 101 is used to obtain a target video to be edited, perform video frame analysis on the target video to be edited, and obtain key video frames;

[0160] A sparse tensor field construction module 102 is configured to generate a sparse tensor field according to the key video frames, and extract global video features of the target video to be edited using the sparse tensor field;

[0161] A topology feature map construction module 103 is configured to construct a topology feature map based on the global video features;

[0162] A cross-frame feature optimization module 104 is configured to perform cross-frame feature optimization on the topological feature map to obtain a depth feature map;

[0163] A dynamic receptive field adjustment module 105 is configured to perform dynamic receptive field adjustment on the depth feature map to obtain a plurality of preliminary clips;

[0164] The clip optimization module 106 is configured to optimize the content of the preliminary clips to obtain a final clip set.

[0165] In one embodiment, when the video key frame extraction module 101 performs video frame analysis on the target video to be edited and obtains key video frames, it is configured to:

[0166] Decoding the target video to be edited to obtain a plurality of initial video frames;

[0167] Performing optical character recognition on the characters in the initial video frame and counting the characters to obtain a total number of the characters;

[0168] Determine whether the total number of characters is greater than a preset threshold;

[0169] If the total number of characters is greater than the number threshold, the target video to be edited is determined to be a text video type;

[0170] Performing valid text recognition on the initial video frame according to the text video type to obtain the number of valid texts in each frame;

[0171] Calculating an average value of all the valid character numbers, and screening out initial video frames corresponding to valid character numbers greater than the average value;

[0172] Using the filtered initial video frame as a key video frame;

[0173] If the total number of characters is less than or equal to the number threshold, the target video to be edited is determined to be an image video type;

[0174] Randomly select two initial video frames with consecutive preset first timestamps as a group of frames to be detected;

[0175] Calculating the image similarity of the frames to be detected according to the image and video type, and screening out a first group of target detection frames whose image similarity is greater than a preset similarity threshold in a preset order;

[0176] An initial video frame with a larger preset first timestamp among the filtered target detection frames is used as a key video frame.

[0177] In one embodiment, when the sparse tensor field construction module 102 generates a sparse tensor field according to the key video frames and extracts global video features of the target video to be edited using the sparse tensor field, it is configured to:

[0178] Obtaining a key video image corresponding to the key video frame, and calculating a gradient magnitude map of the key video image;

[0179] Screening out gradient amplitudes greater than a preset amplitude threshold in the gradient amplitude map, and obtaining a sparse coordinate set corresponding to the screened gradient amplitudes;

[0180] Performing hybrid feature extraction on the sparse coordinate set to obtain a depth feature vector;

[0181] constructing a sparse tensor field according to the depth feature vector;

[0182] The sparse tensor field is aggregated to obtain global video features.

[0183] In one embodiment, when constructing a topology feature map based on the global video features, the topology feature map construction module 103 is configured to:

[0184] Using the sparse coordinate set of each key video frame as a graph node;

[0185] Randomly select two sparse coordinates in the sparse coordinate set as coordinate pairs to be traversed;

[0186] Calculating the two-dimensional space distance of each of the coordinate pairs to be traversed;

[0187] Determining whether the two-dimensional space distance is less than a preset plane distance threshold;

[0188] If the two-dimensional space distance is greater than or equal to the plane distance threshold, no undirected graph edge is established in the coordinate pair to be traversed;

[0189] If the two-dimensional space distance is less than the plane distance threshold, establishing an undirected graph edge in the coordinate pair to be traversed;

[0190] Randomly select two key video frames with consecutive preset second timestamps as adjacent key frames;

[0191] Calculating the three-dimensional space distance between every two sparse coordinates in the adjacent key frames;

[0192] Determining whether the three-dimensional space distance is less than a preset timing threshold;

[0193] If the three-dimensional spatial distance is greater than or equal to the timing threshold, no cross-frame connection edge is established in the adjacent key frames;

[0194] If the three-dimensional spatial distance is less than the temporal threshold, establishing a cross-frame connection edge in the adjacent key frames;

[0195] A topological feature graph is constructed according to the graph nodes, the undirected graph edges, the cross-frame connection edges and the global video features.

[0196] In one embodiment, when performing cross-frame feature optimization on the topological feature map to obtain a depth feature map, the cross-frame feature optimization module 104 is configured to:

[0197] Obtaining an initial attention weight, and using the initial attention weight to perform feature propagation and update on the topological feature map to obtain an updated feature map;

[0198] Calculating an updated attention weight between each pair of adjacent nodes in the updated feature graph;

[0199] The updated attention weight is used as a new initial attention weight, and the initial attention weight is returned to obtain the initial attention weight, and the topological feature map is propagated and updated using the initial attention weight to obtain the step of updating the feature map, and the number of returns is counted;

[0200] When the number of returns is equal to a preset number threshold, the return is ended, and the updated feature map returned for the last time is used as the depth feature map.

[0201] In one embodiment, when the dynamic receptive field adjustment module 105 performs dynamic receptive field adjustment on the depth feature map to obtain a plurality of preliminary clips, it is configured to:

[0202] Performing cluster analysis on the depth feature map to obtain several target clusters;

[0203] Determine the receptive field area according to the cluster center of the target cluster, and calculate the semantic entropy value of the receptive field area;

[0204] Dynamically adjusting the radius of the receptive field area according to the semantic entropy value to obtain an updated cluster;

[0205] Obtaining frame numbers corresponding to the graph nodes in the update cluster in the target video to be edited;

[0206] The smallest frame number is used as the clip start frame, and the largest frame number is used as the clip end frame;

[0207] Obtaining a search frame distance, adjusting boundaries of adjacent frames of the clip start frame according to the search frame distance, and updating the start frame;

[0208] A plurality of initial clipping segments are generated according to the update start frame and the clipping end frame.

[0209] In one embodiment, when performing content optimization on the preliminary clips to obtain a final set of clips, the clip optimization module 106 is configured to:

[0210] Calculating the average semantic entropy of the video frames in the preliminary clip one by one;

[0211] Determining whether the average semantic entropy is greater than a preset average entropy value threshold;

[0212] If the average semantic entropy is less than or equal to the average entropy threshold, deleting the preliminary clipping segments corresponding to the average semantic entropy less than or equal to the average entropy threshold;

[0213] If the average semantic entropy is greater than the average entropy value threshold, the preliminary clipping segment corresponding to the average semantic entropy greater than the average entropy value threshold is used as the clipping retention segment;

[0214] Randomly selecting two clips of the clip with adjacent preset third timestamps as a clip comparison group;

[0215] calculating the time interval and semantic similarity between the comparison groups of the fragments;

[0216] If the time interval is less than a preset time threshold and the semantic similarity is greater than a preset semantic similarity threshold, merging the segment comparison groups to obtain an updated clip;

[0217] The retained clips and the updated clips are aggregated to obtain a final clip set.

[0218] In the present invention, a device for generating clips based on video features is provided. First, the present invention obtains a target video to be edited, performs video frame analysis on the target video to be edited, obtains key video frames, automatically classifies the video into a text video or an image video according to a text quantity threshold, adopts differentiated key frame extraction strategies for different types, effectively improves the accuracy and representativeness of the key video frames, generates a sparse tensor field based on the key video frames, and uses the sparse tensor field to extract the global video features of the target video to be edited, accurately screens out sparse coordinate points representing significant edge and texture features, effectively captures the local details and structural information of the video, and generates global features reflecting the overall content and semantics of the video through aggregation of the sparse tensor field, thereby improving the accuracy of video analysis, and constructs a topological feature map based on the global video features, effectively capturing The spatial structure and temporal relationship of the video content, while integrating global video features, realizes the unified expression of local details and overall semantics of the video, and optimizes the cross-frame features of the topological feature map to obtain a deep feature map, realizes the deep mining and optimization of the spatiotemporal correlation between nodes in the topological feature map, and effectively enhances the fusion and expression capabilities of cross-frame features. Then, the deep feature map is dynamically adjusted to obtain a number of preliminary clips, which can accurately identify the semantically rich and compact content areas in the video, effectively avoiding the information omission or redundancy caused by the fixed range. Finally, the preliminary clips are optimized to obtain the final clip set, and the adjacent clips are merged based on the time interval and semantic similarity to improve the continuity and overall coherence of the clips, which can effectively improve the efficiency of video editing and the accuracy of the generated clip video clips. The specific definition of a clip generation device based on video features can be found in the definition of a clip generation method based on video features above, which will not be repeated here. The various modules in the above-mentioned clip generation device based on video features can be implemented in whole or in part by software, hardware and their combination. The above modules may be embedded in or independent of the processor in the computer device in the form of hardware, or may be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0219] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a method for generating clips based on video features.

[0220] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a method for generating clips based on video features.

[0221] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0222] Obtaining a target video to be edited, performing video frame analysis on the target video to be edited, and obtaining key video frames;

[0223] Generating a sparse tensor field according to the key video frame, and extracting global video features of the target video to be edited using the sparse tensor field;

[0224] Constructing a topological feature map according to the global video features;

[0225] Performing cross-frame feature optimization on the topological feature map to obtain a depth feature map;

[0226] Dynamically adjusting the receptive field of the depth feature map to obtain a number of preliminary clips;

[0227] Content optimization is performed on the preliminary clips to obtain a final clip set.

[0228] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.

[0229] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0230] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0231] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0232] In some implementations of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.

[0233] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can implement:

[0234] Obtaining a target video to be edited, performing video frame analysis on the target video to be edited, and obtaining key video frames;

[0235] Generating a sparse tensor field according to the key video frame, and extracting global video features of the target video to be edited using the sparse tensor field;

[0236] Constructing a topological feature map according to the global video features;

[0237] Performing cross-frame feature optimization on the topological feature map to obtain a depth feature map;

[0238] Dynamically adjusting the receptive field of the depth feature map to obtain a number of preliminary clips;

[0239] Content optimization is performed on the preliminary clips to obtain a final clip set.

[0240] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0241] The computer-readable storage medium may also store at least one computer-executable program / instruction, such as a computer-readable instruction. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.

[0242] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (eg, keyboard, mouse, speaker, etc.).

[0243] The processor can communicate with external devices via an I / O bus via a wired or wireless network.

[0244] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.

[0245] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0246] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0247] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a portion of code, and the above-mentioned module, program segment or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0248] It should be noted that, in this disclosure, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element limited by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0249] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

[0250] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.

Claims

1. A method for generating clips based on video features, characterized in that: The method comprises: Obtaining a target video to be edited, performing video frame analysis on the target video to be edited, and obtaining key video frames; Generating a sparse tensor field according to the key video frame, and extracting global video features of the target video to be edited using the sparse tensor field; Constructing a topological feature map according to the global video features; Performing cross-frame feature optimization on the topological feature map to obtain a depth feature map; Dynamically adjusting the receptive field of the depth feature map to obtain a number of preliminary clips; Content optimization is performed on the preliminary clips to obtain a final clip set.

2. The method for generating clips based on video features according to claim 1, wherein: The step of performing video frame analysis on the target video to be edited to obtain key video frames includes: Decoding the target video to be edited to obtain a plurality of initial video frames; Performing optical character recognition on the characters in the initial video frame and counting the characters to obtain a total number of the characters; Determine whether the total number of characters is greater than a preset threshold; If the total number of characters is greater than the number threshold, the target video to be edited is determined to be a text video type; Performing valid text recognition on the initial video frame according to the text video type to obtain the number of valid texts in each frame; Calculating an average value of all the valid character numbers, and screening out initial video frames corresponding to valid character numbers greater than the average value; Using the filtered initial video frame as a key video frame; If the total number of characters is less than or equal to the number threshold, the target video to be edited is determined to be an image video type; Randomly select two initial video frames with consecutive preset first timestamps as a group of frames to be detected; Calculating the image similarity of the frames to be detected according to the image and video type, and screening out a first group of target detection frames whose image similarity is greater than a preset similarity threshold in a preset order; An initial video frame with a larger preset first timestamp among the filtered target detection frames is used as a key video frame.

3. The method for generating clips based on video features according to claim 1, wherein: Generating a sparse tensor field according to the key video frame and extracting global video features of the target video to be edited using the sparse tensor field includes: Obtaining a key video image corresponding to the key video frame, and calculating a gradient magnitude map of the key video image; Screening out gradient amplitudes greater than a preset amplitude threshold in the gradient amplitude map, and obtaining a sparse coordinate set corresponding to the screened gradient amplitudes; Performing hybrid feature extraction on the sparse coordinate set to obtain a depth feature vector; constructing a sparse tensor field according to the depth feature vector; The sparse tensor field is aggregated to obtain global video features.

4. The method for generating clips based on video features according to claim 3, wherein: The constructing a topological feature map according to the global video features includes: Using the sparse coordinate set of each key video frame as a graph node; Randomly select two sparse coordinates in the sparse coordinate set as coordinate pairs to be traversed; Calculating the two-dimensional space distance of each of the coordinate pairs to be traversed; Determining whether the two-dimensional space distance is less than a preset plane distance threshold; If the two-dimensional space distance is greater than or equal to the plane distance threshold, no undirected graph edge is established in the coordinate pair to be traversed; If the two-dimensional space distance is less than the plane distance threshold, establishing an undirected graph edge in the coordinate pair to be traversed; Randomly select two key video frames with consecutive preset second timestamps as adjacent key frames; Calculating the three-dimensional space distance between every two sparse coordinates in the adjacent key frames; Determining whether the three-dimensional space distance is less than a preset timing threshold; If the three-dimensional spatial distance is greater than or equal to the timing threshold, no cross-frame connection edge is established in the adjacent key frames; If the three-dimensional spatial distance is less than the temporal threshold, establishing a cross-frame connection edge in the adjacent key frames; A topological feature graph is constructed according to the graph nodes, the undirected graph edges, the cross-frame connection edges and the global video features.

5. The method for generating clips based on video features according to claim 1, wherein: The performing cross-frame feature optimization on the topological feature map to obtain a depth feature map includes: Obtaining an initial attention weight, and using the initial attention weight to perform feature propagation and update on the topological feature map to obtain an updated feature map; Calculating an updated attention weight between each pair of adjacent nodes in the updated feature graph; The updated attention weight is used as a new initial attention weight, and the initial attention weight is returned to obtain the initial attention weight, and the topological feature map is propagated and updated using the initial attention weight to obtain the step of updating the feature map, and the number of returns is counted; When the number of returns is equal to a preset number threshold, the return is ended, and the updated feature map returned for the last time is used as the depth feature map.

6. The method for generating clips based on video features according to claim 1, wherein: The dynamic receptive field adjustment is performed on the depth feature map to obtain a plurality of preliminary clips, including: Performing cluster analysis on the depth feature map to obtain several target clusters; Determine the receptive field area according to the cluster center of the target cluster, and calculate the semantic entropy value of the receptive field area; Dynamically adjusting the radius of the receptive field area according to the semantic entropy value to obtain an updated cluster; Obtaining frame numbers corresponding to the graph nodes in the update cluster in the target video to be edited; The smallest frame number is used as the clip start frame, and the largest frame number is used as the clip end frame; Obtaining a search frame distance, adjusting boundaries of adjacent frames of the clip start frame according to the search frame distance, and updating the start frame; A plurality of initial clipping segments are generated according to the update start frame and the clipping end frame.

7. The method for generating clips based on video features according to claim 1, wherein: The step of optimizing the content of the preliminary clips to obtain a final clip set includes: Calculating the average semantic entropy of the video frames in the preliminary clip one by one; Determining whether the average semantic entropy is greater than a preset average entropy value threshold; If the average semantic entropy is less than or equal to the average entropy threshold, deleting the preliminary clipping segments corresponding to the average semantic entropy less than or equal to the average entropy threshold; If the average semantic entropy is greater than the average entropy value threshold, the preliminary clipping segment corresponding to the average semantic entropy greater than the average entropy value threshold is used as the clipping retention segment; Randomly selecting two clips of the clip with adjacent preset third timestamps as a clip comparison group; calculating the time interval and semantic similarity between the comparison groups of the fragments; If the time interval is less than a preset time threshold and the semantic similarity is greater than a preset semantic similarity threshold, merging the segment comparison groups to obtain an updated clip; The retained clips and the updated clips are aggregated to obtain a final clip set.

8. A video feature-based clip generation device, characterized in that: The device comprises: A video key frame extraction module is used to obtain a target video to be edited, perform video frame analysis on the target video to be edited, and obtain key video frames; A sparse tensor field construction module, configured to generate a sparse tensor field according to the key video frames, and extract global video features of the target video to be edited using the sparse tensor field; A topological feature map construction module, configured to construct a topological feature map according to the global video features; A cross-frame feature optimization module is used to perform cross-frame feature optimization on the topological feature map to obtain a depth feature map; A dynamic receptive field adjustment module, configured to dynamically adjust the receptive field of the depth feature map to obtain a plurality of preliminary clips; The clip optimization module is used to optimize the content of the preliminary clips to obtain a final clip set.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating clips based on video features as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating clips based on video features as claimed in any one of claims 1 to 7 is implemented.