Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

588 results about "Video editing" patented technology

Video editing is the manipulation and arrangement of video shots. Video editing is used to structure and present all video information, including films and television shows, video advertisements and video essays. Video editing has been dramatically democratized in recent years by editing software available for personal computers. Editing video can be difficult and tedious, so several technologies have been produced to aid people in this task. Pen based video editing software was developed in order to give people a more intuitive and fast way to edit video.

Video editing method and related equipment

The invention discloses a video editing method and related equipment, and the method comprises the steps: receiving an editing intention of a user, carrying out the semantic analysis of an intention analysis model, and generating an editing copywriting, a style and a structured instruction; the edited copywriting is split according to the edited copywriting; performing cross-modal analysis on video materials by using a video understanding model trained based on an open-source multi-modal model, and accurately positioning sub-lens segments; secondly, editing the sub-lens segments through an editing model to generate a preliminary sheet, performing quantitative scoring through a video scoring model, and obtaining a scoring result according to a multi-dimensional index; and finally, the result is fed back to the editing model, and iterative adjustment is carried out until a final slice is produced. According to the method, the editing efficiency is improved, the manual operation time is shortened, the material adaptability is enhanced, the personalized requirement is met, the editing effect consistency is improved, the video content understanding is deepened, the editing effect is optimized, the technical threshold is reduced, the existing technical problems are solved comprehensively, and an efficient, intelligent and personalized editing scheme is provided.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Ai-enhanced video editing with intermediate data model representation and web-based interface

The present invention relates to a computer-implemented method and system for generating a video editing project using artificial intelligence (AI) and machine learning (ML) techniques. The method includes processing a collection of video clips to generate text-based metadata, receiving selection criteria to identify relevant video clips, and generating a natural language prompt based on the selection criteria. The prompt, comprising instructions and context, is provided to a large language model (LLM), which processes the input and outputs data for constructing a video project data model. The project data model includes timing data for salient snippets within the selected video clips. A dynamic and interactive web-based user interface is rendered to visually represent the project data model, offering a timeline view and editing tools for refining the video project. This system streamlines the video editing process by integrating AI-driven content analysis with user-directed editing, resulting in a tailored video project that aligns with user-defined thematic elements.
Owner:JOBPIXEL INC

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Automatic video editing method based on semantic analysis

The invention relates to an automatic video editing method based on semantic analysis, and belongs to the technical field of video editing. The method comprises the following steps: extracting anchor point information of text content in a streaming media file based on a knowledge graph, establishing a support relationship, and determining a main anchor point, a sub anchor point and a argument anchor point by combining the support relationship and the anchor point information; constructing a dynamic trigger mechanism, setting a verification direction of the dynamic trigger mechanism, determining a semantic extension direction of the anchor point information according to the verification direction, and verifying whether a precondition chain supporting the anchor point information exists or not; and setting a credibility scoring strategy, determining the credibility of the anchor point information according to the logic complexity, the language confidence and the text consistency, forming a sequence set of the anchor point information, and splicing to form a logically coherent video editing text result. According to the method, automatic and intelligent editing of the video is realized based on semantic analysis, and the editing efficiency and quality are remarkably improved.
Owner:BEIJING GONGXIN INTERNET TECHNOLOGY CO LTD

Video intelligent self-adaptive editing method and system based on deep learning

The invention provides an intelligent self-adaptive video editing method and system based on deep learning, and relates to the technical field of video processing.The method comprises the steps that firstly, a semantic mapping relation between a to-be-edited video material and a preset editing requirement is established, and an editing requirement mapping result is generated, the preset editing demand comprises a content style and a rhythm control demand, and then semantic feature association processing is carried out based on the mapping result to obtain a semantic association feature set comprising lens unit content semantic features and rhythm association features; then calling a pre-trained editing decision model (including a semantic matching module and a rhythm adjusting module) to carry out editing strategy matching on the set, generating a preliminary editing strategy set, generating an initial video editing scheme according to the preliminary editing strategy set, carrying out parameter adjustment on the initial scheme according to a strategy optimization suggestion output by the model, and carrying out video editing on the initial scheme; and a final video editing scheme is obtained, and intelligent self-adaptive editing of the video is realized.
Owner:WEIMAI TECH CO LTD

Video editing and artistic creation system based on AI intelligence

The invention discloses a video editing and artistic creation system based on AI intelligence, which relates to the technical field of video editing and comprises a data acquisition module, a feature analysis module, a feature library construction module, a rule matching module and a split mirror generation module, the data acquisition module is used for acquiring traditional opera stage video data and corresponding audience feedback audio waveform data, and the feature analysis module is used for extracting spatial-temporal feature vectors of stylized actions from the traditional opera stage video data. The feature library construction module is used for constructing an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors; the method has the beneficial effects that dynamic association analysis is performed by constructing the action feature library and combining audience feedback audio data, the short video sequence conforming to the art rule is automatically generated, and the method has the advantage of improving the Chinese opera performance video editing efficiency and the art expressive force.
Owner:PACO VIDEO TECH (HANGZHOU) CO LTD

Video editing method based on grid layout alternate diffusion and multi-attention control

The invention relates to the technical field of video analysis, in particular to a video editing method based on grid layout alternate diffusion and multi-attention control, and the method comprises the steps: segmenting an original video frame sequence into a plurality of grids, each grid comprising a plurality of pixel space video frames which are continuously arranged, and forming grid data; mapping the gridding data to a low-dimensional submerged space through an encoder, and generating initial submerged space feature data; the initial submerged space feature data are edited, the editing process comprises a diffusion process and a sampling process, the diffusion process is based on a pre-trained stable diffusion model, and a time attention module is embedded in the diffusion process; in the sampling process, executing an odd-even time step alternate replacement strategy on the grid layout to promote cross-grid global consistency, and dynamically fusing attention maps of a reconstruction branch and an editing branch according to a timestamp threshold to generate de-noised data; and decoding, splitting and recombining the de-noised data through a decoder to generate an edited continuous video frame sequence.
Owner:ANHUI UNIV

Intelligent video editing method fusing human face features and human voice features

The invention relates to an intelligent video editing method fusing human face features and human voice features, which comprises the steps of input preprocessing, multi-modal analysis, fusion scoring and automatic editing, and adopts a multi-modal mode for analysis, so that the recall rate of key segments is improved, and the efficiency of editing is improved. A personalized editing strategy is supported, facial expression changes and voice emotion peak values are aligned through dynamic time warping, a CLIP-like structure is used for training a bimodal encoder, the human face and the voice are mapped to a unified vector space, the association weight of the human face and the voice is automatically learned, and the intelligent degree of editing is improved.
Owner:BEIJING HEJUHUITONG E-COMMERCE CO LTD

Temporally consistent and semantics guided text-based video editing generative artificial intelligence (AI) model with improved initialization

A processor-implemented method performed for text-based video editing includes receiving a video input and a text prompt. The video input includes a sequence of video frames. Features of the video input are extracted to generate a latent representation of the video input. Noise is injected to the latent representation of the video input to generate a noise injected latent. The noise is conditioned on the video input. An artificial neural network (ANN) model processes the noise injected latent based on the text prompt to adapt the video input according to the text prompt.
Owner:QUALCOMM INC

Video editing method and device, electronic equipment and nonvolatile storage medium

The invention discloses a video editing method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring an audio data stream of a target video, and segmenting the audio data stream into a plurality of audio clips; determining a classification result corresponding to the audio clip, and determining a time period corresponding to the audio clip as a candidate ball hitting time period when the classification result is that the audio clip contains the ball hitting sound; acquiring a video frame corresponding to the candidate ball hitting time period in the target video, and judging that the candidate ball hitting time period is a real ball hitting time period under the condition that the visual feature of the video frame is matched with a preset ball hitting rule; and determining an editing time point according to the real ball hitting time period, and editing the target video according to the editing time point to obtain a ball hitting round video clip. The technical problem that the accuracy of ball hitting detection segment detection is low due to the fact that ball hitting detection is conducted only through sound and is easily interfered by environmental noise in the prior art is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Multi-modal portrait video editing method, electronic device, and storage medium

The present invention provides a multi-modal portrait video editing method, which can be applied to the technical field of video editing. The method comprises: given a portrait video, preprocessing same to obtain camera parameters, a human identity coefficient, a human expression coefficient, a human pose coefficient, a human semantic segmentation map, and a two-dimensional portrait mask; on the basis of a neural Gaussian texture mechanism, embedding a learnable three-dimensional Gaussian feature into a parameterized human geometric surface, using a neural renderer to convert a three-dimensional Gaussian splatting feature map into an image, and optimizing the reconstruction of a three-dimensional portrait on the basis of RGB and segmentation map information of the video; and using an iterative dataset update technique to distill knowledge of a multi-modal two-dimensional image generation model into three-dimensional portrait editing, and using expression similarity guidance and a face-aware portrait editing model to improve editing quality. The method elevates a two-dimensional editing task to three-dimensional space, and ensures good three-dimensional consistency and temporal consistency. By means of the knowledge of the multi-modal generation model, high-quality portrait video editing functions can be achieved.
Owner:UNIV OF SCI & TECH OF CHINA

Audio data selection for video matching using generative artificial intelligence model

A video editing system leverages a generative artificial intelligence (AI) model to identify songs to overlay on a video. The video editing system extracts a set of key frames from the video and prompts the generative AI model to generate a video narrative for the video. A video narrative is a text description of the plot, theme, feel, or other characteristics of the video. The video editing system uses the video narrative to prompt the generative AI model again to generate a set of descriptor tags for the video based on the video narrative. Descriptor tags are strings that represent themes, features, or characteristics of the song. The video editing system uses an audio tagging system to score a set of songs based on the set of descriptor tags and presents a selected subset of the set of songs based on the scores of the songs.
Owner:BEACON STREET TECHNOLOGIES LLC

Short play video editing method, system and device based on multiple modes and medium

The invention discloses a multi-mode-based short episode video editing method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing of an original short episode video, and obtaining a second short episode video, subtitles and a subtitle timestamp; analyzing the subtitles, and generating a short episode abstract according to an analysis result of the subtitles; according to the short episode abstract and the highlight editing prompt, performing plot analysis and editing on the second short episode video, and performing editing to form a first video set; scoring videos in the first video set according to a multi-dimensional scoring rule, and screening out M highlight video clips before scoring; and adjusting the timestamps of the M highlight video clips according to the subtitle timestamps, and outputting the adjusted M highlight video clips. According to the method and the device, the full link from short video input to short play mixed video output is constructed, a user only needs to input the to-be-edited short play video, the mixed video can be automatically generated, and the video editing efficiency is improved while the video editing threshold is reduced.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Video editing method and device, electronic equipment and storage medium

The invention discloses a video editing method and device, electronic equipment and a storage medium. The method comprises the following steps: segmenting a to-be-edited video to obtain a plurality of video clips; understanding text generation and subtitle generation are carried out on each video clip, and a clip understanding text and a clip subtitle of each video clip are obtained; determining a target description text based on fragment understanding texts and / or fragment subtitles corresponding to a plurality of target video fragments selected from the plurality of video fragments; and performing video synthesis based on the plurality of target video clips and the target description text to obtain a target video corresponding to the to-be-edited video. According to the method provided by the invention, the video editing efficiency is relatively high, and the video editing cost is relatively low.
Owner:SHENZHEN SIYUAN ELECTRONICS TECH CO LTD

Visual and text search interface for text-based video editing

Embodiments of the present invention provide systems, methods, and computer storage media for a visual and text search interface used to navigate a video transcript. In an example embodiment, a freeform text query triggers a visual search for frames of a loaded video that match the freeform text query (e.g., frame embeddings that match a corresponding embedding of the freeform query), and triggers a text search for matching words from a corresponding transcript or from tags of detected features from the loaded video. Visual search results are displayed (e.g., in a row of tiles that can be scrolled to the left and right), and textual search results are displayed (e.g., in a row of tiles that can be scrolled up and down). Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.
Owner:ADOBE INC

Online video instance segmentation method and system based on mask propagation

The invention provides an online video instance segmentation method and system based on mask propagation, and the method comprises the following steps: S1, achieving the cross-frame propagation of a target mask through a mask propagation model, obtaining a prediction mask of a current frame, and achieving the correlation of a target through the calculation of the intersection-union ratio of the prediction mask; and S2, performing back propagation on the target by using the mask propagation model so as to complement the missed target mask and enhance the continuity of the track. Objects between different frames are associated on the basis of a mask propagation mechanism and in combination with the intersection-to-union ratio of masks, and meanwhile, in order to relieve the scene that a segmentation model fails to segment shielded objects, non-significant objects and the like, a reverse mask propagation technology is provided to complement the failed frames. The method aims at performing fine segmentation and continuous tracking on various scenes such as intelligent monitoring, automatic driving, augmented reality and video editing on the target in the video, and has wide research and application prospects.
Owner:FUDAN UNIVERSITY

Video editing model based on common editing of text and image and construction method thereof

The invention provides a video editing model based on text and image common editing and a construction method thereof. The video editing model introduces an optical flow guide mask fusion module and a multi-modal feature recognition and segmentation module into a denoising diffusion implicit model; the method comprises the following steps: inputting an original video into a denoising diffusion implicit model to carry out forward diffusion noise addition to obtain a multi-frame submerged space generation frame, and inputting the multi-frame submerged space generation frame into an optical flow guide mask fusion module to carry out inter-frame feature alignment to obtain a time consistency submerged space generation frame; a text prompt, an image prompt and an original video are input into a multi-modal feature recognition and segmentation module to be aligned and positioned to obtain a condition vector, the condition vector and a time consistency submerged space generation frame are subjected to iterative denoising to generate a target editing video, and dynamic feature modulation is performed in each denoising process; a video editing model based on text and image common editing efficiently edits a video under combined guidance of text and image prompts.
Owner:HANGZHOU GISWAY INFORMATION TECH CO LTD

Diffusion model video local editing method and system based on mask guidance

The invention provides a diffusion model video local editing method and system based on mask guidance. The method comprises the following steps: encoding a video sequence of a target video to obtain a first video frame submerged space feature; performing coarse-grained mask labeling on a target editing area in the target video to obtain mask information; determining a space attention weight according to the first video frame potential space feature; determining a second video frame potential space feature according to the first video frame potential space feature, the mask information and the space attention weight; determining a time attention weight according to the second video frame potential space feature; determining a third video frame potential space feature according to the mask information, the second video frame potential space feature and the time attention weight; and decoding the hidden space feature of the third video frame to generate an edited video. According to the method, frame-by-frame accurate masking is not needed, the time-space consistency of local editing of the video can be effectively enhanced, an unedited area is kept unchanged, and a high-quality and stable video editing effect is achieved.
Owner:HEFEI UNIV OF TECH

Video processor, method, apparatus, storage medium, and program product

The embodiment of the invention provides a video processor, a video processing method, video processing equipment, a storage medium and a program product. In the scheme, after the display layer receives the special effect style configuration, the algorithm engine layer calls the matched physical mathematical model according to the target special effect style and the basic video, outputs the high-fidelity special effect description file, and improves the special effect parameter accuracy; the rendering adaptation layer dynamically selects an optimal rendering engine according to the resources of all the rendering engines, and the time consumption of single-example rendering is shortened; two-layer hierarchical decoupling is adopted, so that description file generation and rendering execution are parallel, batch description files can be shared by multiple preview positions after being prepared at one time, rendering resources are only input when real preview requirements are met, and idling is avoided; the display layer, the algorithm engine layer, the rendering adaptation layer and the rendering layer are sequentially relayed to form a compact link of quasi-generation-fast rendering-playback, so that waiting and repeated export links of a user in a subsequent video editing stage are remarkably reduced, and the video editing efficiency is integrally improved.
Owner:BEIJING 58 INFORMATION TTECH CO LTD

Three-dimensional scene video editing method based on point cloud guidance

The invention provides a three-dimensional scene video editing method based on point cloud guidance, and the method comprises the steps: obtaining an original video of a three-dimensional scene, and estimating the three-dimensional point cloud of the scene in a specified frame of the video and camera parameters of each video frame; determining an editing reference image of the specified frame according to the image of the specified frame, the pixel-level mask and the editing area description text; estimating the edited depth of the specified frame according to the edited reference image to obtain an edited three-dimensional point cloud corresponding to the specified frame; according to the mask of the specified frame and the pre-edit depth map and the post-edit depth map corresponding to the image of the specified frame, constructing a three-dimensional grid model used for surrounding an edit area, and transmitting the mask of the specified frame to the view angle of other frames by using the three-dimensional grid model to obtain masks of other frames; and obtaining a point cloud rendering image of each frame rendered according to the edited three-dimensional point cloud and the camera parameters of each frame, generating an image editing result of each frame according to the point cloud rendering image of each frame, the image, the mask and the editing reference image, and splicing the image editing result into an edited video.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Automatic video editing method and device based on text and shot similarity, and terminal

The invention discloses an automatic video editing method and device based on text and shot similarity and a terminal, and belongs to the technical field of artificial intelligence. The method comprises the following steps: determining the end time of a mixed video based on a music interval identification result of specified music; selecting a lens of which the main body label is scenery as a lens head lens; performing similarity analysis with other lenses based on a subject recognition result, a behavior recognition result and a motion calculation result, calculating visual similarity between the lenses based on a deep learning model, recognizing repeated or redundant pictures, and selecting a high-energy lens as an in-chip high-combustion lens; retrieving and intercepting a corresponding lens as a tail lens based on a corresponding segment of the selected end word; and carrying out audio and video mixed cutting and assembling on the specified music, the head lens, the in-chip high-combustion lens and the tail lens. According to the invention, by fusing multi-modal feature analysis and intelligent editing logic generation, efficient and high-quality video automatic production is realized.
Owner:SHENZHEN KUKAI SOFTWARE TECH CO LTD

Video editing processing method and system based on artificial intelligence

The invention discloses a video editing processing method and system based on artificial intelligence, and the method comprises the steps: obtaining an initial video, extracting the visual semantic features of the initial video through a visual semantic model, and extracting the sound semantic features of the initial video through a sound semantic model; an indication text input by a user is obtained, a preference model of the user is obtained, and the preference model is constructed based on historical behavior data of the user; and processing the visual semantic features and the sound semantic features based on the indication text and the preference model to obtain an edited target video. In this way, the video can be efficiently and accurately edited, and it is ensured that the video editing result can meet the personalized requirements of the user.
Owner:SHANGHAI YOUANMI INFORMATION TECHNOLOGY CO LTD

Automatic video editing method and system for clip list recommendation

The invention discloses an automatic video editing method and system for film list recommendation, and the method comprises the steps: obtaining film data, and carrying out the preprocessing of the film data; according to the preprocessed film data, generating a recommendation list and basic data of a film list by utilizing a preset large model and cue words; generating a copywriting corresponding to the film list data according to the recommendation list and the basic data, and selecting corresponding film resources and background music; and according to the copywriting, the film resource and the background music, performing video synthesis editing, and outputting an automatic edited video recommended by a film list. According to the method, the material combinatorial logic can be dynamically adjusted according to the sheet theme, deep mining and structured application are performed on the sheet recommendation data, and high-efficiency and large-scale sheet recommendation video generation is realized.
Owner:SHENZHEN KUKAI SOFTWARE TECH CO LTD

CPU idle power distribution system for PC based on AI

The invention relates to the technical field of electric power distribution, and provides an AI-based CPU idle electric power distribution system for a PC, which can flexibly optimize CPU electric power distribution according to real-time data of different scenes by virtue of an AI algorithm, meets diversified workload requirements, has the advantages of energy conservation, performance and adaptability, brings more efficient and stable use experience for PC users, and has better application prospects in the aspect of energy conservation. The system can accurately sense load changes and adjust the power supply voltage and frequency of a CPU in real time, in office, game and video editing scenes, power consumption is reduced compared with a traditional system, electric power waste is effectively reduced, the use cost is reduced, in terms of performance, the system greatly improves user experience, software and a browser run smoothly during office, the frame rate is stable during game playing, and the service life of the system is prolonged. Video editing and rendering are fast, and preview is not blocked.
Owner:SHANGHAI YINGZHONG INFORMATION TECH CO LTD

Video intelligent editing method, device and equipment based on graph-to-text large model

The invention relates to the technical field of large models, solves the problems of high cost, low efficiency and the like due to the fact that video editing depends on manual processing in the prior art, and provides an intelligent video editing method, device and equipment based on a graph-to-text large model. The method comprises the following steps: preprocessing an original video to be edited, and obtaining a plurality of frames of first video images after preprocessing; performing clustering analysis on the first video images to obtain a plurality of image sets; inputting each image set into a pre-trained graph-to-text large model to obtain a key tag corresponding to each image set; and comparing each key tag with a preset user instruction, and according to a comparison result, determining a target image set in each image set as an edited target video clip. According to the method, intelligent identification and editing of the video content are realized through a large model, the manual workload is reduced, and reliable technical support is provided for video editing and processing.
Owner:NINGBO SIMSHINE INTELLIGENT TECH CO LTD

Audio and video processing method, system and equipment based on large model and medium

The invention discloses an audio and video processing method, system and device based on a large model, and a medium. The method comprises the following steps: acquiring an audio and video to be processed and a demand processing instruction; obtaining a processing instruction set of the audio and video editing tool, and performing knowledge base construction processing on the processing instruction set according to the large model to obtain an instruction knowledge base; performing demand identification processing on the demand processing instruction according to the instruction knowledge base to obtain an operation instruction link; performing material content selection and instruction parameter calculation processing on the operation instruction link to obtain a target processing instruction; and performing generation processing on the to-be-processed audio and video according to the target processing instruction to obtain a target audio and video. The embodiment of the invention can improve the audio and video processing efficiency, and can be widely applied to the technical field of audio and video processing.
Owner:IMUSIC CULTURE & TECH CO LTD

Video and audio integrated packaging and editing method, device and equipment

The invention relates to the technical field of multimedia processing, and discloses a video and audio integrated packaging and editing method, device and equipment, and the method comprises the steps: disassembling original video and audio data, respectively extracting audio and video information streams, carrying out semantic analysis through a pre-trained video and audio information analysis model, and generating a semantic sequence of the audio and the video; based on the sequences, key events are identified, an editing positioning shaft is generated, finally, video and audio editing and format packaging are carried out according to the positioning shaft, the requirements of a multi-end platform are met, and corresponding editing files are generated, according to the method, through intelligent analysis and automatic editing, the editing efficiency is greatly improved, manual intervention is reduced, and the editing efficiency is improved. Moreover, the method can accurately adapt to the demands of different platforms, improves the availability and propagation effect of video and audio contents, and solves a problem that the video and audio editing efficiency is low in the prior art.
Owner:BEIJING FILM & VIDEO ORIGIN FILM & TELEVISION CULTURE MEDIA CO LTD

Video editing method and device based on artificial intelligence, equipment and medium

The invention relates to the technical field of artificial intelligence and image processing, can be applied to the field of intelligent medical treatment and finance, and discloses a video editing method and device based on artificial intelligence, equipment and a medium, and the method comprises the steps: obtaining a to-be-edited video, and carrying out the preprocessing of the to-be-edited video; based on a gating mechanism, according to the video content, carrying out dynamic division of spatial-temporal characteristics on the characteristic pattern of the video frame sequence in the video to be edited to obtain a plurality of content-coherent spatial-temporal partitions, and fusing the characteristics of the plurality of spatial-temporal partitions to obtain a fused characteristic; analyzing the fusion features based on a pre-trained segmentation model, determining editing boundary points, and editing the to-be-edited video according to the editing boundary points to generate at least one target video; and generating a final edited video of the to-be-edited video according to the target video. According to the method, feature regions can be dynamically divided according to video contents based on a gating mechanism, space-time joint self-adaptive partitioning is realized, and redundant calculation can be avoided.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video editing method and device based on multi-modal large model, equipment and medium

The invention relates to the technical field of video editing, solves the problem that how to obtain a video clip is not involved in the prior art, and provides a video editing method, device and equipment based on a multi-modal large model and a medium, and the method comprises the steps: obtaining image description information input by a user and a to-be-edited video; performing semantic segmentation on the to-be-edited video according to the image description information to obtain a plurality of initial segmentation fragments in the to-be-edited video; for each initial boundary frame in each initial segmentation segment, if an adjacent video frame which is adjacent to the initial boundary frame and is not included in the initial segmentation segment exists in the video to be edited, obtaining a target boundary frame according to the adjacent video frame and the initial boundary frame; and determining a target video clip in the video to be edited according to the two target boundary frames in the initial segmentation clip. Dynamic adjustment and optimization of the boundary can be realized, so that the target video clip in the video can be accurately determined.
Owner:NINGBO SIMSHINE INTELLIGENT TECH CO LTD