Video editing method and device based on artificial intelligence, equipment and medium
Through dynamic division of spatiotemporal features based on the gating mechanism and lightweight neural network processing, the problem of redundancy in video clip computing in the existing technology is solved, and efficient calculation and precise resource allocation of video clips are realized.
Patent Information
- Application Number
- CN202510840316.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-22
AI Technical Summary
The existing neural network model calculates redundancy problems during video clipping, especially uniform processing of video frames, which ignores the heterogeneity of spatiotemporal features and leads to wasted computing resources.
The feature map of the video frame sequence is dynamically divided into spatial and temporal features based on the gating mechanism, and the spatial and temporal partitions that are related to the content are generated, and the partition features are processed and fused in parallel through a lightweight neural network, and the clip boundary points are determined in combination with the pre-trained segmentation model to generate the target video.
The calculation efficiency of video clips is optimized, redundant calculations of simple backgrounds or static areas are avoided, and accurate allocation of computing resources and efficient processing of video clips are realized.
Smart Images

Figure CN120529151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and image processing technology, and in particular to an artificial intelligence-based video editing method, device, equipment and medium. Background Art
[0002] With the development of the self-media era, video editing plays a very important role today. It not only helps to enrich self-media resources, but also helps to improve user viewing efficiency. Traditional video editing methods mainly rely on manual annotation or rule-based feature extraction, which requires video editors to consume a lot of energy and time, and is inefficient. Especially when the videos to be edited are medical imaging videos, medical-related teaching videos, or promotional videos related to financial products, editors with relevant medical or financial product knowledge must be familiar with the video content before they can perform reasonable editing and annotation, which reduces the efficiency of video editing.
[0003] In recent years, with the development of deep learning technology, more and more users have applied AI to automatically edit videos. For example, user-applied AI built based on a 3D convolutional neural network model predicts editing points based on the temporal modeling of the 3D convolutional neural network and the video content understanding of the attention mechanism, thereby editing and generating short videos. However, this method uniformly processes video frames when extracting features from the video, ignoring the heterogeneity of spatiotemporal features, resulting in a large amount of computation wasted on simple backgrounds or static areas, causing computational redundancy problems. Summary of the Invention
[0004] The present invention provides a video editing method, device, equipment and medium based on artificial intelligence to solve the technical problem of computational redundancy when users use AI to perform video editing based on existing neural network models.
[0005] In a first aspect, a video editing method based on artificial intelligence is provided, comprising:
[0006] Obtaining a video to be edited, and preprocessing the video to be edited;
[0007] Dynamically dividing the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on the gating mechanism to obtain a plurality of spatiotemporal partitions with coherent contents, and fusing the features of the plurality of spatiotemporal partitions to obtain a fused feature;
[0008] Analyzing fusion features based on a pre-trained segmentation model, determining clipping boundary points, and clipping the video to be clipped according to the clipping boundary points to generate at least one target video;
[0009] A final edited video of the video to be edited is generated according to the target video.
[0010] In a second aspect, a video editing device based on artificial intelligence is provided, comprising:
[0011] A preprocessing module is used to obtain the video to be edited and preprocess the video to be edited;
[0012] A feature dynamic segmentation module is used to dynamically segment the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on a gating mechanism to obtain a plurality of spatiotemporal partitions with coherent content, and fuse the features of the plurality of spatiotemporal partitions to obtain a fused feature;
[0013] A segmentation module, configured to analyze fusion features based on a pre-trained segmentation model, determine clipping boundary points, and clip the video to be clipped according to the clipping boundary points to generate at least one target video;
[0014] A post-processing module is used to generate a final edited video of the video to be edited based on the target video.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned video editing method when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned video editing method are implemented.
[0017] In the scheme implemented by the above-mentioned artificial intelligence-based video editing method, device, equipment and medium, the video to be edited can be obtained through the client and pre-processed; and based on the gating mechanism, the feature map of the video frame sequence in the video to be edited is dynamically divided into temporal and spatial features according to the video content to obtain multiple content-coherent temporal and spatial partitions, and the features of the multiple temporal and spatial partitions are fused to obtain fused features; then the fused features are analyzed based on the pre-trained segmentation model to determine the editing boundary points, and the video to be edited is edited according to the editing boundary points to generate at least one target video; finally, the final edited video of the video to be edited is generated according to the target video. In the present invention, when editing medical imaging videos, teaching videos and financial technology promotion or teaching videos for smart medical care, the feature map of the video frame sequence in the video to be edited can be dynamically divided into temporal and spatial features based on the content-aware gating mechanism, that is, the feature area is dynamically divided according to the video content, thereby realizing the spatiotemporal joint adaptive partitioning of the video feature map, avoiding redundant calculations of simple backgrounds or static areas, realizing dynamic adjustment of computing resource allocation, and optimizing the computing efficiency of video editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 is a schematic diagram of an application environment of a video editing method according to an embodiment of the present invention;
[0020] Figure 2 is a flow chart of a video editing method according to an embodiment of the present invention;
[0021] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S20;
[0022] Figure 4 is a structural diagram of a video editing device according to an embodiment of the present invention;
[0023] Figure 5 yes Figure 4 A structural diagram of a specific implementation of the dynamic feature division module;
[0024] Figure 6 is a structural diagram of a computer device in one embodiment of the present invention;
[0025] Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] The video editing method based on artificial intelligence provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can obtain the video to be edited through the client and pre-process the video to be edited; and based on the gating mechanism, the feature map of the video frame sequence in the video to be edited is dynamically divided into spatiotemporal features according to the video content to obtain multiple content-coherent spatiotemporal partitions, and the features of the multiple spatiotemporal partitions are fused to obtain fused features; then the fused features are analyzed based on the pre-trained segmentation model to determine the editing boundary points, and the video to be edited is edited according to the editing boundary points to generate at least one target video; finally, the final edited video of the video to be edited is generated according to the target video. In the present invention, when editing medical imaging videos, teaching videos, and promotion or teaching videos of financial technology for smart medical care, the feature map of the video frame sequence in the video to be edited can be dynamically divided into spatiotemporal features based on the content-aware gating mechanism, that is, the feature area is dynamically divided according to the video content, and the spatiotemporal joint adaptive partitioning of the video feature map is realized, which can avoid redundant calculation of simple background or static areas, realize dynamic adjustment of computing resource allocation, and optimize the computing efficiency of video editing while ensuring processing accuracy. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0028] See also Figure 2 As shown, Figure 2 A flowchart of a video editing method based on artificial intelligence provided by an embodiment of the present invention includes the following steps S10-S40:
[0029] S10: Obtain the video to be edited, and pre-process the video to be edited.
[0030] In the present invention, the video to be edited can be a long video with a duration of more than 5 minutes, or a short video with a duration of 1-5 minutes. For example, it can be an image video of a surgical procedure or a medical-related teaching video, a promotional video of a medical product, a promotional or teaching video of a financial technology product, or a sports game video, etc.
[0031] In this step, the pre-processing of the video to be edited specifically includes: standardizing the resolution and color of the video to be edited to unify the resolution, color, and brightness range of the video frames in the video to be edited.
[0032] S20: Based on the gating mechanism, dynamically divide the spatiotemporal features of the feature map of the video frame sequence in the video to be edited according to the video content to obtain multiple content-coherent spatiotemporal partitions, and fuse the features of the multiple spatiotemporal partitions to obtain fused features.
[0033] In this step, a gating mechanism is used to dynamically focus on important features in different spatiotemporal regions of the video to achieve the division of content-coherent regions.
[0034] Specifically, if Figure 3 As shown, step S20 includes the following steps S21-S25:
[0035] S21: Extracting spatiotemporal features of a video frame sequence in the video to be edited, and generating a feature map of the video frame sequence.
[0036] In this step, the backbone feature extractor is used to extract the spatiotemporal features of the video frame sequence in the video to be edited, and the feature map F∈R of the video frame sequence is obtained. T×H×W×C , where T represents the length of the time dimension, H and W represent the spatial height and width respectively, and C represents the number of feature channels. Preferably, the backbone feature extractor can be composed of a backbone model such as SwinTransformer-3D or VisionTransformer.
[0037] S22: Dynamically partitioning the spatiotemporal features of the feature map of the video frame sequence according to the video content of the video to be edited based on a gating mechanism to generate a plurality of content-aware partition masks.
[0038] Specifically, this step is: according to the video content of the video to be edited, use the gate mechanism formula M i =σ(W g *F+b g ) Dynamically partition the spatiotemporal features of the feature map of the video frame sequence to generate multiple content-aware partition masks M i ; Where F represents the feature map, F∈R T×H×W×C ,σ represents the sigmoid activation function, W g and b g Represent the learnable convolution kernel parameters and bias terms, * represents the convolution operation, M i Represents the partition mask of the i-th partition. The value range of each mask matrix is between 0 and 1, that is, M i ∈[0,1] T×H×W , which represents the probability that the corresponding spatiotemporal position belongs to the partition.
[0039] S23: Obtain multiple content-coherent spatiotemporal partitions according to the multiple content-aware partition masks.
[0040] In this step, the formula S is used based on multiple content-aware partition masks. i =F⊙M i Get multiple content-coherent spatiotemporal partitions; where ⊙ represents element-by-element multiplication, S i Represents a time-space partition. Each time-space partition Si Both retain the parts of the feature map that are related to specific content, such as moving objects, static backgrounds, or specific areas.
[0041] For steps S21-S23, the correlation of time and space dimensions is simultaneously considered through the gating mechanism, and the feature map is dynamically divided into multiple content-coherent areas (i.e., time-space partitions), realizing joint time-space partitioning, which can adapt to content changes (such as object movement, camera switching, etc.), intuitively display the model's focus area (such as medical surgery teaching videos can focus on the instructor's hand movements and surgical sites), suppress irrelevant areas (such as static backgrounds), avoid the limitations of fixed partitions, improve computing efficiency, and facilitate long video processing; for example, when editing medical surgery teaching videos, based on this step of the present invention, the feature map of the video frame can be divided into four main time-space partitions: time-space partition 1: surgical site (high temporal dynamics); time-space partition 2: surgeon's hand movements (medium time-space dynamics); time-space partition 3: movements of other doctors (such as assistants) (low dynamics but critical); and time-space partition 4: background areas such as the operating table (low dynamics but secondary). Moreover, the partitioning process is completely data-driven and can adapt to different video contents. For example, it can also be used to edit sports game videos. During editing, this step based on the present invention can also divide the feature map of the video frame into four main spatiotemporal partitions: spatiotemporal partition 1: football motion trajectory; spatiotemporal partition 2: main player actions; spatiotemporal partition 3: goal area; and spatiotemporal partition 4: audience area.
[0042] S24: Based on multiple lightweight neural networks, the features of the multiple spatiotemporal partitions are processed in parallel.
[0043] In this embodiment, the lightweight neural network adopts the EfficientNet-Lite architecture. For the i-th partition feature S i , the processing process is: i =φ i (S i ), where φ i represents the i-th lightweight neural network, v i is the processed eigenvector, v i ∈R D , D is the feature dimension.
[0044] In this embodiment, the EfficientNet-Lite architecture is used to reduce the number of parameters and enhance feature expression capabilities. It is understandable that each lightweight neural network can optimize for specific types of content within a spatiotemporal partition when processing it, and the corresponding number of temporal convolutional layers can vary. For example, a lightweight neural network processing highly spatiotemporal dynamics can include more temporal convolutional layers, enabling precise allocation of computing resources.
[0045] S25: Based on the attention mechanism, the features of each spatiotemporal partition after parallel processing are fused to obtain fused features.
[0046] In this step, weighted fusion is performed based on the attention weights of different spatiotemporal partitions. Specifically, the attention weights of different partitions are first calculated: Where q∈R D is a learnable query vector, K=[v1,…,v N ] T is the feature v of each spatiotemporal partition after parallel processing i The matrix composed of is the key vector matrix, is a scaling factor used to stabilize training, D is the dimension of the K vector, and N is the number of spatiotemporal partitions; then weighted fusion is performed: Get the fusion feature z.
[0047] In the present invention, the fused feature z contains the most relevant information of each spatiotemporal partition, and automatically adjusts the contribution of different spatiotemporal partitions through the attention mechanism. At the same time, it realizes the feature fusion of content context awareness and maintains the semantic association between spatiotemporal partitions.
[0048] S30: Analyze fusion features based on a pre-trained segmentation model, determine clipping boundary points, and clip the video to be clipped according to the clipping boundary points to generate at least one target video.
[0049] In this step, the fusion features are analyzed based on a pre-trained segmentation model constructed by a temporal convolutional network to obtain the boundary probability of each time point to determine the editing boundary points. That is, the temporal convolutional network processes the input fusion features, extracts the temporal features of the fusion features, and outputs the boundary probability of the features corresponding to each time point through a sigmoid function. If the boundary probability is greater than the preset probability threshold, this time point is the editing boundary point (for example, the 5th and 6th minutes in medical and financial teaching videos). The video to be edited can then be edited according to the editing boundary points, accurately editing at least one target video and avoiding interference from irrelevant content.
[0050] S40: Generate a final edited video of the video to be edited according to the target video.
[0051] In the present invention, target videos are optimized and / or merged to generate a final edited video of the video to be edited. Specifically, in this embodiment, the target videos are optimized to filter out conflicting target videos, and when there are multiple target videos, the similarity between adjacent target videos is calculated, and adjacent target videos with similarities greater than a preset similarity threshold are merged to generate the final edited video of the video to be edited. When there is only one target video, this target video serves as the final edited video of the video to be edited.
[0052] It can be seen that in the above scheme, when editing medical imaging videos, teaching videos, and promotion or teaching videos of financial technology for smart medical care, the feature maps of the video frame sequence in the edited video can be dynamically divided into spatiotemporal features based on the content-aware gating mechanism, that is, the feature areas are dynamically divided according to the video content, thereby realizing the spatiotemporal joint adaptive partitioning of the video feature map, avoiding redundant calculations of simple backgrounds or static areas, and realizing dynamic adjustment of computing resource allocation. While ensuring processing accuracy, the computational efficiency of video editing is optimized, and different lightweight neural networks are used to process different spatiotemporal partitions in parallel. Each lightweight neural network can optimize the corresponding content in the spatiotemporal partition, realizing accurate allocation of computing resources, and maintaining the semantic association between different spatiotemporal partitions through the calculation of cross-partition attention weights and the fusion of each spatiotemporal partition, thereby ensuring the coherence and naturalness of the final edited video, and significantly improving the editing accuracy and efficiency in complex scenes (such as multi-target concurrent actions).
[0053] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0054] In one embodiment, a video editing device based on artificial intelligence is provided, which corresponds to the video editing method based on artificial intelligence in the above embodiment. Figure 4 As shown, the video editing device includes a pre-processing module 110, a feature dynamic division module 120, a segmentation module 130 and a post-processing module 140. The functional modules are described in detail as follows:
[0055] The pre-processing module 110 is used to obtain the video to be edited and pre-process the video to be edited;
[0056] Feature dynamic division module 120, such as Figure 5 As shown, it includes a dynamic partitioning unit 121 and a feature processing unit 122, wherein the dynamic partitioning unit 121 is used to dynamically divide the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on the gating mechanism to obtain multiple spatiotemporal partitions with coherent content; the feature processing unit 122 is used to fuse the features of the multiple spatiotemporal partitions to obtain a fused feature;
[0057] The segmentation module 130 is configured to analyze the fusion features based on the pre-trained segmentation model, determine the clipping boundary points, and clip the video to be clipped according to the clipping boundary points to generate at least one target video;
[0058] A post-processing module 140 is configured to generate a final edited video of the video to be edited based on the target video;
[0059] In one embodiment, if Figure 5 As shown, the dynamic partitioning unit 121 includes a feature extraction unit 1211 and a gated partitioning unit 1212, wherein:
[0060] The feature extraction unit 1211 is used to: extract the spatiotemporal features of the video frame sequence in the video to be edited, and generate a feature map of the video frame sequence;
[0061] The gated partitioning unit 1212 is used to: dynamically divide the spatiotemporal features of the feature map of the video frame sequence according to the video content of the video to be edited based on the gating mechanism, generate multiple content-aware partitioning masks; and obtain multiple content-coherent spatiotemporal partitions based on the multiple content-aware partitioning masks.
[0062] In one embodiment, the gated partition unit 1212 is specifically configured to:
[0063] According to the video content of the video to be edited, the gate mechanism formula M is used i =σ(W g *F+b g ) Dynamically partition the spatiotemporal features of the feature map of the video frame sequence to generate multiple content-aware partition masks M i Among them, M i represents the partition mask of the i-th partition, F represents the feature map, F∈R T×H×W×C ,σ represents the sigmoid activation function, W g and b g They represent the learnable convolution kernel parameters and bias terms respectively, and * represents the convolution operation.
[0064] In one embodiment, the gated partition unit 1212 is further configured to:
[0065] Based on multiple content-aware partition masks using formula S i =F⊙M i Get multiple content-coherent spatiotemporal partitions; where ⊙ represents element-by-element multiplication, S i Representing spatiotemporal partitions.
[0066] In one embodiment, if Figure 5 As shown, the feature processing unit 122 includes a feature parallel processing unit 1221 and a fusion unit 1222, wherein:
[0067] The feature parallel processing unit 1221 is used to: perform parallel processing on the features of the multiple spatiotemporal partitions based on multiple lightweight neural networks;
[0068] The fusion unit 1222 is used to fuse the features of each spatiotemporal partition after parallel processing based on the attention mechanism to obtain a fused feature.
[0069] It can be understood that the gated partitioning unit 1212 and the feature processing unit 122 in the video editing device of the present invention can be seamlessly integrated into the existing video editing processing flow, that is, embedded in the existing video editing processing model, for example, inserted after the backbone extraction network, without changing the overall framework, and end-to-end video editing optimization can be achieved. At the same time, the modular design supports incremental unit updates without the need for system-level retraining.
[0070] In one embodiment, the segmentation module 130 is specifically configured to:
[0071] Based on the pre-trained segmentation model constructed by the temporal convolutional network, the fusion features are analyzed to obtain the boundary probability of each time point to determine the clip boundary points.
[0072] In one embodiment, the post-processing module 140 is specifically configured to:
[0073] The target video is optimized, and when there are multiple target videos, the similarity between adjacent target videos is calculated, and the adjacent target videos whose similarity is greater than a preset similarity threshold are merged to generate a final edited video of the video to be edited.
[0074] The present invention provides a video editing device, which obtains a video to be edited and pre-processes the video to be edited; and based on a gating mechanism, dynamically divides the spatiotemporal features of a feature map of a video frame sequence in the video to be edited according to the video content to obtain a plurality of content-coherent spatiotemporal partitions, and fuses the features of the plurality of spatiotemporal partitions to obtain a fusion feature; then analyzes the fusion feature based on a pre-trained segmentation model, determines the editing boundary points, and edits the video to be edited according to the editing boundary points to generate at least one target video; finally, generates a final edited video of the video to be edited according to the target video, that is, when working, the spatiotemporal features of the feature map of the video frame sequence in the video to be edited can be dynamically divided based on the content-aware gating mechanism, that is, the feature area is dynamically divided according to the video content, thereby realizing the spatiotemporal joint adaptive partitioning of the video feature map, avoiding redundant calculation of simple background or static areas, realizing dynamic adjustment of computing resource allocation, and optimizing the computing efficiency of video editing while ensuring processing accuracy.
[0075] The specific definition of the video editing device can be found in the definition of the video editing method above and will not be repeated here. The various modules in the above-mentioned video editing device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0076] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a video editing method based on artificial intelligence.
[0077] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a video editing method based on artificial intelligence
[0078] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0079] Obtaining a video to be edited, and preprocessing the video to be edited;
[0080] Dynamically dividing the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on the gating mechanism to obtain a plurality of spatiotemporal partitions with coherent contents, and fusing the features of the plurality of spatiotemporal partitions to obtain a fused feature;
[0081] Analyzing fusion features based on a pre-trained segmentation model, determining clipping boundary points, and clipping the video to be clipped according to the clipping boundary points to generate at least one target video;
[0082] A final edited video of the video to be edited is generated according to the target video.
[0083] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0084] Obtaining a video to be edited, and preprocessing the video to be edited;
[0085] Dynamically dividing the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on the gating mechanism to obtain a plurality of spatiotemporal partitions with coherent contents, and fusing the features of the plurality of spatiotemporal partitions to obtain a fused feature;
[0086] Analyzing fusion features based on a pre-trained segmentation model, determining clipping boundary points, and clipping the video to be clipped according to the clipping boundary points to generate at least one target video;
[0087] A final edited video of the video to be edited is generated according to the target video.
[0088] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0089] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0090] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0091] The above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention and are therefore intended to be included within the scope of protection of the present invention. Furthermore, any software tools or components not owned by the Company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A video editing method based on artificial intelligence, characterized in that: include: Obtaining a video to be edited, and preprocessing the video to be edited; Dynamically dividing the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on the gating mechanism to obtain a plurality of spatiotemporal partitions with coherent contents, and fusing the features of the plurality of spatiotemporal partitions to obtain a fused feature; Analyzing fusion features based on a pre-trained segmentation model, determining clipping boundary points, and clipping the video to be clipped according to the clipping boundary points to generate at least one target video; A final editing video of the video to be edited is generated according to the target video.
2. The video editing method based on artificial intelligence according to claim 1, characterized in that: The gating mechanism is used to dynamically divide the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content to obtain a plurality of content-coherent spatiotemporal partitions, including: Extracting spatiotemporal features of a video frame sequence in the video to be edited, and generating a feature map of the video frame sequence; Dynamically partitioning the spatiotemporal features of the feature map of the video frame sequence according to the video content of the video to be edited based on a gating mechanism to generate a plurality of content-aware partition masks; A plurality of content-coherent spatiotemporal partitions are obtained according to a plurality of content-aware partition masks.
3. The video editing method based on artificial intelligence according to claim 2, characterized in that: The gating mechanism dynamically divides the spatiotemporal features of the feature map of the video frame sequence according to the video content of the video to be edited, and generates multiple content-aware partition masks, including: According to the video content of the video to be edited, the gate mechanism formula M is used i =σ(W g *F+b g ) Dynamically partition the spatiotemporal features of the feature map of the video frame sequence to generate multiple content-aware partition masks M i Among them, M i represents the partition mask of the i-th partition, F represents the feature map, F∈R T×H×W×C ,σ represents the sigmoid activation function, W g and b g They represent the learnable convolution kernel parameters and bias terms respectively, and * represents the convolution operation.
4. The video editing method based on artificial intelligence according to claim 3, characterized in that: The obtaining of a plurality of content-coherent spatiotemporal partitions according to a plurality of content-aware partition masks comprises: Based on multiple content-aware partition masks using formula S i =F⊙M i Get multiple content-coherent spatiotemporal partitions; where ⊙ represents element-by-element multiplication, S i Representing spatiotemporal partitions.
5. The video editing method based on artificial intelligence according to claim 1, wherein: The fusing of the features of the plurality of spatiotemporal partitions to obtain a fused feature includes: Based on multiple lightweight neural networks, the features of multiple spatiotemporal partitions are processed in parallel; Based on the attention mechanism, the features of each spatiotemporal partition after parallel processing are fused to obtain the fused features.
6. The video editing method based on artificial intelligence according to claim 1, characterized in that: The pre-trained segmentation model is used to analyze the fusion features and determine the clip boundary points, including: Based on the pre-trained segmentation model constructed by the temporal convolutional network, the fusion features are analyzed to obtain the boundary probability of each time point to determine the clip boundary points.
7. The video editing method based on artificial intelligence according to claim 1, characterized in that: Generating a final edited video of the video to be edited according to the target video includes: The target video is optimized, and when there are multiple target videos, the similarity between adjacent target videos is calculated, and the adjacent target videos whose similarity is greater than a preset similarity threshold are merged to generate a final edited video of the video to be edited.
8. A video editing device based on artificial intelligence, characterized in that: include: A preprocessing module is used to obtain the video to be edited and preprocess the video to be edited; A feature dynamic segmentation module is used to dynamically segment the spatiotemporal features of the feature graph of the video frame sequence in the video to be edited according to the video content based on a gating mechanism to obtain a plurality of spatiotemporal partitions with coherent content, and fuse the features of the plurality of spatiotemporal partitions to obtain a fused feature; A segmentation module, configured to analyze fusion features based on a pre-trained segmentation model, determine clipping boundary points, and clip the video to be clipped according to the clipping boundary points to generate at least one target video; A post-processing module is used to generate a final edited video of the video to be edited based on the target video.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the video editing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the video editing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video editing method based on artificial intelligence
CN112218005A
Video editing method and device thereof, electronic equipment and storage medium
CN113676671A
Video abstract algorithm and system based on gated multi-head position attention mechanism
CN115002559A
Video analysis method and system of monitoring terminal
CN118781526A
Cited By
Method and device for intelligently repairing oral error in video, storage medium and program product
CN121908089A