Method and device for editing video based on text

By using a multimodal alignment model to calculate the similarity between text and image feature vectors during video editing, determining the target image frame and performing video interception, the problems of low editing efficiency, poor graphics and text consistency and limited scalability in the prior art are solved, and efficient and accurate video editing and batch processing capabilities are achieved.

CN120201243AInactive Publication Date: 2025-06-24RUISHI DATA TECH (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510669641.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing methods based on text-based video editing have problems such as low editing efficiency, poor graphics and text consistency and limited scalability, which are difficult to adapt to the production needs of automated content.

Method used

By obtaining the to-process video and the corresponding commentary text, the image frame is extracted according to the preset frame interval and the image feature vector is determined, the target commentary word is selected from the commentary text in combination with the natural paragraph order and its text feature vector is determined. The similarity between the two is calculated using the multimodal alignment model to determine the target image frame, and the video is intercepted based on the timing position of the image frame to generate the target clip video.

Benefits of technology

It improves video editing efficiency, ensures consistency of graphics and text, solves the problem of front and rear camera splitting caused by time estimation deviation or lens omission in manual interception, and adapts to the large-scale processing needs of multiple videos and massive commentary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201243A_ABST
    Figure CN120201243A_ABST
Patent Text Reader

Abstract

The invention provides a method for editing a video based on a text, and the method comprises the steps: obtaining a to-be-processed video and a corresponding explanation text, extracting an image frame from the to-be-processed video according to a preset frame number interval, and determining an image feature vector corresponding to the image frame; wherein the number of the extracted image frames is at least one; selecting an unprocessed target commentary from the commentary text according to a natural paragraph sequence, and determining a text feature vector corresponding to the target commentary; determining a target image frame having the highest similarity with the target commentary vector according to the text feature vector and the image feature vector; and performing video interception on the to-be-processed video based on the time sequence position corresponding to the target image frame to generate a target edited video. Accurate association of commentary and image pictures is realized through feature matching, and manual frame-by-frame processing is avoided; and video interception is carried out based on the target image frame, so that narrative integrity and cross-fragment continuity are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a method and device for text-based video clipping. Background Art

[0002] Currently, short videos of movie explanations are very popular. The production of short videos of movie explanations generally adopts a manual clipping process. The editor needs to watch the original film completely to understand the plot context, then write the explanation text according to the story, and manually locate the picture segments in the film that match each sentence of the text one by one. After extracting the target segments through a clipping tool, all segments are then spliced into a complete video in the order of the text.

[0003] However, the existing methods for text-based video clipping generally have problems such as low clipping efficiency and poor text-image consistency. Specifically, matching text with image frames sentence by sentence requires repeatedly playing back and searching the film content, which is extremely time-consuming. For example, when processing a two-hour movie, locating the picture corresponding to a single sentence of the explanation may take dozens of minutes. Secondly, the editor needs to be highly familiar with the details of the film. Otherwise, it is difficult to quickly associate the explanation with the picture, resulting in low efficiency and high mis-matching rate when processing new films or cross-genre clipping. In addition, manual operations are prone to problems such as deviation in the time of segment extraction, omission of key shots, or repeated extraction. Finally, this method is difficult to be extended to batch processing of multiple films or series of works and cannot adapt to the requirements of automated content production. Summary of the Invention

[0004] In view of the above problems, this application is proposed to provide a method for text-based video clipping that overcomes or at least partially solves the above problems, including: A method for text-based video clipping, including the steps of: Obtaining a video to be processed and the corresponding explanation text, extracting image frames from the video to be processed at a preset frame number interval, and determining the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; Selecting an unprocessed target explanation word from the explanation text according to the natural paragraph order, and determining the text feature vector corresponding to the target explanation word; Determining the target image frame with the highest similarity to the target explanation word vector according to the text feature vector and the image feature vector; Performing video clipping on the video to be processed based on the temporal position corresponding to the target image frame to generate a target clipped video.

[0005] Preferably, the step of extracting image frames from the video to be processed at a preset frame number interval and determining the image feature vectors corresponding to the image frames includes: Extract image frames from the video to be processed according to a preset frame number interval, and construct an image database; Input each image frame in the image database into a multi-modal alignment model, and calculate the image feature vector corresponding to each image frame.

[0006] Preferably, the step of selecting an unprocessed target commentary from the commentary text according to the natural paragraph order and determining the text feature vector corresponding to the target commentary includes: Segment the commentary text into multiple commentaries according to semantic integrity, and construct a commentary processing sequence; wherein, each commentary is provided with a processing identifier, and the processing identifier includes an unprocessed identifier and a processed identifier; Select the first commentary with an unprocessed identifier from the processing sequence as the target commentary according to the original text arrangement order of the commentary text; Input the target commentary into a multi-modal alignment model, and calculate the text feature vector corresponding to the target commentary.

[0007] Preferably, it further includes: Mark the status identifier of the target commentary for which the target clip video has been generated in the processing sequence as processed.

[0008] Preferably, the step of determining the target image frame with the highest similarity to the target commentary vector based on the text feature vector and the image feature vector includes: Input the text feature vector and the image feature vector into a multi-modal alignment model, and calculate the similarity between the text feature vector of the target commentary and the image feature vector of each image frame; Sort the vector similarities in descending order, and select the image frame with the highest similarity to the target commentary vector as the target image frame according to the sorting result.

[0009] Preferably, the step of performing video interception on the video to be processed based on the corresponding temporal position of the target image frame to generate a target clip video includes: Obtain the first temporal position of the target image frame in the video to be processed; Set the target image frame as the temporal midpoint of the target clip video; Intercept a video segment with a preset duration from the video to be processed as the target clip video according to the first temporal position and the temporal midpoint.

[0010] Preferably, the step of intercepting a video segment with a preset duration from the video to be processed as the target clip video according to the first temporal position and the temporal midpoint includes: Calculate the start time point and end time point of the target clipped video based on the timing midpoint and the preset duration; Convert the start time point and the end time point into the start clipping timing position and the end clipping timing position in the video to be processed; Clip a video segment from the video to be processed as the target clipped video according to the start clipping timing position and the end clipping timing position.

[0011] Preferably, the step of performing video clipping on the video to be processed based on the timing position of the target image frame to generate a video segment to be processed includes: Obtain the second timing position of the target image frame in the video to be processed; Obtain a transition frame corresponding to the target image frame from the video to be processed, and obtain the timing position corresponding to the transition frame; Clip a video segment from the video to be processed as the target clipped video according to the second timing position and the timing position corresponding to the transition frame.

[0012] Preferably, the step of clipping a video segment from the video to be processed as the target clipped video according to the second timing position and the timing position of the transition frame includes: Calculate the start transition frame and the end transition frame with the closest timing distance to the target image frame according to the second timing position and the timing position corresponding to the transition frame respectively; Clip a video segment from the video to be processed as the target clipped video according to the timing position of the start transition frame and the timing position of the end transition frame.

[0013] A device for clipping a video based on text, comprising: An image feature extraction module, configured to extract image frames from a video to be processed according to a preset frame number interval, and determine an image feature vector corresponding to the image frame; A text feature extraction module, configured to select an unprocessed target commentary from the commentary text according to the natural paragraph order, and determine a text feature vector corresponding to the target commentary; A cross-modal matching module, configured to determine a target image frame with the highest similarity to the target commentary vector according to the text feature vector and the image feature vector; A video generation module, configured to perform video clipping on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video.

[0014] This application has the following advantages: In an embodiment of the present application, in view of the problems of "low efficiency, poor graphic-text consistency, and limited scalability" in the prior art, the present application provides a solution for intercepting a target video through feature extraction and matching and automatic expansion, specifically as follows: obtaining a video to be processed and a corresponding commentary text, extracting image frames from the video to be processed at a preset frame number interval, and determining an image feature vector corresponding to the image frame; wherein, the number of the extracted image frames is at least one frame; selecting an unprocessed target commentary word from the commentary text according to the natural paragraph order, and determining a text feature vector corresponding to the target commentary word; determining a target image frame with the highest similarity to the target commentary word vector according to the text feature vector and the image feature vector; and performing video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video. By calculating the matching similarity of the image feature vector and the text feature vector, and using the image frame with the highest similarity as the target image frame, it is not necessary to repeatedly watch the video to be processed manually and screen the clips frame by frame, thereby shortening the traditional clip processing time, improving the video clip efficiency, and at the same time ensuring that the scene described in the commentary word is accurately corresponding to the target image frame, avoiding graphic-text deviation. Through video expansion based on the timing position corresponding to the target image frame, it is ensured that the intercepted target clipped video has both independent narrative integrity and cross-segment coherence, and solves the problem of the disconnection of the front and rear shots caused by time estimation deviation or shot omission in manual interception. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the description of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 FIG. is a schematic flowchart of a method for clipping a video based on text provided by an embodiment of the present application; Figure 2 FIG. is a schematic module structure diagram of a device for clipping a video based on text provided by an embodiment of the present application; Figure 3 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, features, and advantages of this application more apparent and understandable, the following provides a more detailed description of this application in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0018] The inventors found through analyzing the prior art that: existing video editing methods rely on a linear operation process dominated by manual experience. During the editing process, manual frame-by-frame search is time-consuming, and traditional keyword matching is difficult to associate the commentary with deep semantics such as continuous actions and scene context, resulting in a high false matching rate. At the same time, the interception boundaries rely on subjective estimation, which easily leads to segment interception deviation or omission of key shots. Therefore, inventing a method for editing videos based on text to solve the problems of low efficiency, deviation in text-image logic, fragmentation between front and rear shots, and adapting to the large-scale processing requirements of multiple videos and a vast amount of commentary has become a technical problem to be urgently solved.

[0019] Refer to Figure 1 , which shows a method for editing videos based on text provided by an embodiment of this application; The method includes: S110. Obtain the video to be processed and the corresponding commentary text, extract image frames from the video to be processed at preset frame intervals, and determine the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; S120. Select an unprocessed target commentary from the commentary text according to the natural paragraph order, and determine the text feature vector corresponding to the target commentary; S130. Determine the target image frame with the highest similarity to the target commentary vector according to the text feature vector and the image feature vector; S140. Perform video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target edited video.

[0020] In the embodiments of the present application, in view of the problems of "low efficiency, poor graphic-text consistency, and limited scalability" in the prior art, the present application provides a solution for intercepting a target video through feature extraction and matching and automatic expansion, specifically as follows: obtaining a video to be processed and its corresponding commentary text, extracting image frames from the video to be processed according to a preset frame number interval, and determining the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; selecting an unprocessed target commentary word from the commentary text according to the natural paragraph order, and determining the text feature vector corresponding to the target commentary word; determining the target image frame with the highest similarity to the target commentary word vector according to the text feature vector and the image feature vector; and performing video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video. By calculating the matching similarity between the image feature vector of the image frame and the text feature vector of the target commentary word, and using the image frame with the highest similarity as the target image frame, there is no need for manual repeated viewing of the video to be processed and frame-by-frame screening and editing, thereby shortening the traditional editing processing time, improving the video editing efficiency, and at the same time ensuring that the scene described in the commentary word corresponds precisely to the target image frame, avoiding graphic-text deviation. Through video expansion based on the timing position corresponding to the target image frame, it is ensured that the generated target clipped video has both independent narrative integrity and cross-segment coherence, solving the problem of fragmentation of front and back shots caused by time estimation deviation or shot omission in manual interception.

[0021] Next, a method for clipping a video based on text in this exemplary embodiment will be further described.

[0022] As described in step S110 above, obtain a video to be processed and its corresponding commentary text, extract image frames from the video to be processed according to a preset frame number interval, and determine the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame.

[0023] In an embodiment of the present invention, the specific process of "extracting image frames from the video to be processed according to a preset frame number interval and determining the image feature vectors corresponding to the image frames" described in step S110 can be further described in combination with the following description.

[0024] As described in the following steps, extract image frames from the video to be processed according to a preset frame number interval and construct an image database; As described in the following steps, input each frame of the image frames in the image database into a multi-modal alignment model, and calculate the image feature vectors corresponding to each frame of the image frames.

[0025] It should be noted that the video to be processed refers to the original video material without editing, which includes continuous video frames and associated audio content; the category of the video in this embodiment is not restricted and can cover various types such as movies, short videos, documentaries, live recordings, etc.; the explanatory text refers to a set of structured text descriptions associated with the content of the video to be processed, which is composed of several independent and semantically complete explanatory words arranged in the order of natural paragraphs, and each explanatory word corresponds to a complete description of an independent event, action, or scene in the video. The value of the preset frame number interval is dynamically adjusted according to the video content redundancy. For conventional videos, frames are extracted at a fixed frame rate by default. For high-dynamic videos, the frame interval is shortened in the action-intensive segments and lengthened in the static or gentle segments to avoid interference from redundant frames and prevent omission of image frames.

[0026] It should be noted that the multi-modal alignment model is a pre-trained text-image alignment model. This model is trained on a text-image dataset through contrastive learning. Its text encoder and image encoder map natural language and images to the same vector space respectively, so as to realize the unified processing of multi-modal tasks.

[0027] In a specific embodiment, multiple image frames are extracted according to the preset frame number interval, and the extracted image frames are stored in the image database in the order of timestamps. The pre-trained multi-modal alignment model is called to extract features from each image frame in the image database, and after being calculated by the model, an image feature vector is output.

[0028] As described in the above step S120, an unprocessed target explanatory word is selected from the explanatory text according to the order of natural paragraphs, and the text feature vector corresponding to the target explanatory word is determined.

[0029] In an embodiment of the present invention, the specific process of "selecting an unprocessed target explanatory word from the explanatory text according to the order of natural paragraphs and determining the text feature vector corresponding to the target explanatory word" described in step S120 can be further described in combination with the following description.

[0030] As described in the following steps, the explanatory text is segmented into multiple explanatory words according to semantic integrity, and an explanatory word processing sequence is constructed; wherein, each of the explanatory words has a processing identifier, and the processing identifier includes an unprocessed identifier and a processed identifier; As described in the following steps, the first explanatory word with an unprocessed identifier is selected from the processing sequence as the target explanatory word according to the original text arrangement order of the explanatory text; As described in the following steps, the target explanatory word is input into the multi-modal alignment model, and the text feature vector corresponding to the target explanatory word is calculated.

[0031] In an embodiment of the present invention, it further includes: marking the status identifier of the target commentary corresponding to the target clip video that has been generated in the processing sequence as processed.

[0032] It should be noted that semantic integrity means that a single commentary needs to express an independent event, action, or scene description. For example, according to punctuation rules, the commentary text is segmented into sentences or semantic groups, or adjacent sentences describe the same continuous action, then they are merged into a single commentary to ensure the integrity of the action. The natural paragraph order refers to the original text arrangement order of the commentary script, which usually corresponds to the timeline of the video content or the logical development of the event.

[0033] The processing identifier is used to achieve the sequential processing and status tracking of the commentaries to avoid duplicate processing. For example, if there are commentary 1, commentary 2, commentary 3..., when the system processes commentary 2, the status identifier of commentary 1 is marked as processed, and commentary 3 and subsequent commentaries remain unprocessed. After the feature extraction and subsequent similarity matching of commentary 2 are completed, its status identifier is updated to processed, and the first unprocessed commentary 3 in the processing sequence automatically becomes the object to be processed until the status of all commentaries in the processing sequence is marked as processed.

[0034] In a specific embodiment, it is segmented into multiple commentaries based on the semantic integrity rule, a status identifier is added to each commentary, and a processing sequence is constructed. The first unprocessed commentary in the sequence is selected as the current target commentary according to the original text arrangement order of the commentary script, and this commentary is locked to prevent concurrent processing conflicts. The target commentary is input into the multi-modal alignment model to calculate a text feature vector with the same dimension as the image feature vector.

[0035] As described in step S130 above, the target image frame with the highest similarity to the target commentary vector is determined based on the text feature vector and the image feature vector.

[0036] In an embodiment of the present invention, the specific process of "determining the target image frame with the highest similarity to the target commentary vector based on the text feature vector and the image feature vector" described in step S130 can be further described in combination with the following description.

[0037] As described in the following steps, the text feature vector and the image feature vector are input into the multi-modal alignment model to calculate the similarity between the text feature vector of the target commentary and the image feature vector of each frame of the image; As described in the following steps, the vector similarities are sorted in descending order, and the image frame with the highest similarity to the target commentary vector is selected as the target image frame according to the sorting result.

[0038] It should be noted that the similarity between the text and the image feature vectors is measured by the Euclidean distance, Manhattan distance or cosine distance. In this embodiment, the cosine similarity is used for measurement. In the vector space model, it determines the similarity degree by calculating the cosine value of the included angle between two vectors. For two non-zero vectors A and B, the value range of the cosine similarity is [-1, 1]. When the cosine similarity is 1, it means that the two vectors are completely similar; when it is -1, it means they are completely opposite; when it is 0, it means that the two vectors are orthogonal, that is, perpendicular to each other and have no similar components.

[0039] If there are multiple image frames with the same similarity to the target commentary and all are the highest values, the candidate frame closest to the temporal position of the previous processed target image frame in terms of distance is preferentially selected to maintain the temporal coherence of the clipped video; if the highest similarity is lower than the preset threshold, it is determined that there is no matching picture, and manual review or automatic association with the context commentary for re-matching is triggered.

[0040] In a specific embodiment, the text feature vector of the input target commentary and the image feature vector of the image frame are input; the cosine similarity between the text feature vector and each image feature vector is calculated in sequence; the cosine similarities are sorted in descending order, and the image frame ranked first is selected as the target image frame. As described in step S140 above, the video to be processed is intercepted based on the temporal position corresponding to the target image frame to generate a target clipped video.

[0041] In an embodiment of the present invention, the specific process of "intercepting the video to be processed based on the temporal position corresponding to the target image frame to generate a target clipped video" described in step S140 can be further described in combination with the following description.

[0042] As described in the following steps, obtain the first temporal position of the target image frame in the video to be processed; As described in the following steps, set the target image frame as the temporal midpoint of the target clipped video; As described in the following steps, according to the first temporal position and the temporal midpoint, intercept a video segment with a preset duration from the video to be processed as the target clipped video.

[0043] It should be noted that taking the target image frame as the midpoint of the intercepted segment ensures that this frame is in the core position in the clipped video, so as to retain the picture content associated before and after it. For example, if the target frame is the starting frame of an action, the interception range needs to cover the complete evolution process of the action from preparation to completion, rather than just intercepting an isolated picture.

[0044] In an embodiment of the present invention, the specific process of "intercepting a video segment with a preset duration from the video to be processed as the target clip video according to the first timing position and the timing midpoint" can be further described in combination with the following description.

[0045] As described in the following steps, calculate the start time point and end time point of the target clip video according to the timing midpoint and the preset duration; As described in the following steps, convert the start time point and the end time point into the start intercept timing position and the end intercept timing position in the video to be processed; As described in the following steps, intercept a video segment from the video to be processed as the target clip video according to the start intercept timing position and the end intercept timing position.

[0046] It should be noted that the preset duration is defined according to the video type. For example, the intercepted duration of long videos is longer to retain the narrative logic, and that of short videos is shorter to focus on the highlight segments. If the calculated start intercept timing position is earlier than the start point of the video or the end intercept timing position is later than the end point of the video, the out-of-bounds part will be automatically truncated to ensure the coherence of the finished video. If a scene transition point, such as a shot transition or a background mutation, is detected within the intercepted range by the scene transition detection tool, the end point of the intercept will be corrected to the nearest scene boundary to avoid content fragmentation.

[0047] In a specific embodiment, obtain the first timing position of the target image frame in the video to be processed and set it as the timing midpoint of the target clip video; calculate the start timing position and the end intercept timing position according to the timing midpoint, the preset duration set according to the video type, and intercept the corresponding video segment from the video to be processed as the target clip video. If there is an overlap between the segment and the generated clip video, only the non-overlapping part is retained, and finally, the target clip video corresponding to the target commentary is output. As described in step S140 above, perform video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clip video.

[0048] In an embodiment of the present invention, the specific process of "performing video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clip video" described in step S140 can be further described in combination with the following description.

[0049] As described in the following steps, obtain the second timing position of the target image frame in the video to be processed; As described in the following steps, obtain the transition frame corresponding to the target image frame from the video to be processed and obtain the timing position corresponding to the transition frame; As described in the following steps, a video segment is intercepted from the video to be processed as the target clipped video according to the second timing position and the timing position corresponding to the transition frame.

[0050] In an embodiment of the present invention, the specific process of "intercepting a video segment from the video to be processed as the target clipped video according to the second timing position and the timing position corresponding to the transition frame" can be further described in combination with the following description.

[0051] As described in the following steps, the starting transition frame and the ending transition frame with the closest timing distance to the target image frame are respectively calculated according to the second timing position and the timing position corresponding to the transition frame; As described in the following steps, a video segment is intercepted from the video to be processed as the target clipped video according to the timing position of the starting transition frame and the timing position of the ending transition frame.

[0052] It should be noted that a transition refers to the transition point from one scene to another in a video, manifested as changes in visual or audio signals such as shot switching, sudden picture changes, and black field transitions, which are used to separate different narrative units; the starting transition frame is the starting boundary frame of the scene where the target image frame is located, and the ending transition frame is the ending boundary frame of the scene where the target image frame is located. By selecting the starting and ending transition frames closest to the target image frame, it is ensured that the intercepted range only includes the complete content of the scene where the target frame is located, avoiding logical breaks caused by cross-scene interception. In this embodiment, a scene transition detection tool such as Scene Detect is used to detect the transition frame corresponding to the target image frame in the video to be processed.

[0053] In a specific embodiment, the second timing position of the target image frame in the video to be processed is obtained, and a scene transition detection tool is used to obtain the transition frame corresponding to the target image frame from the video to be processed, and the timing position corresponding to the transition frame in the video to be processed is obtained to establish a transition frame list, which includes the timing positions of transition frame 1, transition frame 2, transition frame 3... transition frame N. In the transition frame list, all transition frames earlier than the timing position of the target image frame are filtered, and the frame with the largest timing position, that is, the frame closest to the timing position of the target frame, is selected as the starting transition frame; at the same time, all transition frames later than the timing position of the target image frame are filtered, and the frame with the smallest timing position, that is, the frame closest to the target frame, is selected as the ending transition frame. A video segment is intercepted from the video to be processed through the timing positions of the starting transition frame and the ending transition frame to generate the target clipped video. In an embodiment of the present application, it further includes: saving the target clipped video to the local temporary queue.

[0054] It should be noted that for each target clipped video corresponding to a single commentary, it is stored in the local temporary queue according to the natural paragraph order of its corresponding commentary, ensuring that no secondary sorting is required during subsequent splicing. When the status flags of all commentaries are marked as processed, the videos in the temporary queue are seamlessly merged into a complete commentary video through a video splicing tool.

[0055] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, refer to the corresponding descriptions in the method embodiments.

[0056] Refer to Figure 2 , which shows a device for text-based video clipping provided by an embodiment of the present application; Specifically, it includes: An image feature extraction module 210, configured to extract image frames from a video to be processed according to a preset frame number interval, and determine an image feature vector corresponding to the image frames; A text feature extraction module 220, configured to select an unprocessed target commentary from the commentary text according to the natural paragraph order, and determine a text feature vector corresponding to the target commentary; A cross-modal matching module 230, configured to determine a target image frame with the highest similarity to the target commentary vector according to the text feature vector and the image feature vector; A video generation module 240, configured to perform video clipping on the video to be processed based on the timing position corresponding to the target image frame, and generate a target clipped video.

[0057] In an embodiment of the present invention, the image feature extraction module 210 includes: A database construction sub-module, configured to extract image frames from a video to be processed according to a preset frame number interval, and construct an image database; An image feature calculation sub-module, configured to input each image frame in the image database into the multi-modal alignment model, and calculate an image feature vector corresponding to each image frame.

[0058] In an embodiment of the present invention, the text feature extraction module 220 includes: A sequence construction sub-module, configured to segment the commentary text into multiple commentaries according to semantic integrity, and construct a commentary processing sequence; wherein each commentary is provided with a processing flag, and the processing flag includes an unprocessed flag and a processed flag; A target commentary determination sub-module, configured to select the first commentary with an unprocessed flag from the processing sequence as the target commentary according to the original text arrangement order of the commentary text; A text feature calculation sub-module, configured to input the target commentary into the multi-modal alignment model and calculate a text feature vector corresponding to the target commentary.

[0059] In an embodiment of the present invention, it further includes: a status flag update sub-module, configured to mark the status flag of the target commentary of the target clip video that has been generated in the processing sequence as processed.

[0060] In an embodiment of the present invention, the cross-modal matching module 230 includes: A similarity calculation sub-module, configured to input the text feature vector and the image feature vector into the multi-modal alignment model and calculate the similarity between the text feature vector of the target commentary and the image feature vector of each image frame; A target image frame determination sub-module, configured to perform a descending order sorting on the vector similarities and select, according to the sorting result, the image frame with the highest similarity to the target commentary vector as the target image frame.

[0061] In an embodiment of the present invention, the video generation module 240 includes: A timing position acquisition sub-module, configured to acquire a first timing position of the target image frame in the video to be processed; A timing midpoint determination sub-module, configured to set the target image frame as the timing midpoint of the target clip video; A video clip sub-module, configured to intercept a video segment of a preset duration from the video to be processed as the target clip video according to the first timing position and the timing midpoint.

[0062] In an embodiment of the present invention, the video clip sub-module includes: A time boundary determination unit, configured to calculate a start time point and an end time point of the target clip video according to the timing midpoint and the preset duration; A timing position conversion unit, configured to convert the start time point and the end time point into a start intercept timing position and an end intercept timing position in the video to be processed; A video segment interception unit, configured to intercept a video segment from the video to be processed as the target clip video according to the start intercept timing position and the end intercept timing position.

[0063] In another embodiment of the present invention, the video generation module 240 includes: A timing position acquisition sub-module, configured to acquire a second timing position of the target image frame in the video to be processed; A transition boundary determination sub-module, configured to obtain a transition frame corresponding to the target image frame from the video to be processed, and obtain the timing position corresponding to the transition frame; A video clip sub-module, configured to intercept a video segment from the video to be processed as the target clipped video according to the second timing position and the timing position corresponding to the transition frame.

[0064] In an embodiment of the present invention, the video clip sub-module includes: A start and end transition frame determination unit, configured to respectively calculate a start transition frame and an end transition frame with the closest timing distance to the target image frame according to the second timing position and the timing position corresponding to the transition frame; A video segment interception unit, configured to intercept a video segment from the video to be processed as the target clipped video according to the timing position of the start transition frame and the timing position of the end transition frame.

[0065] Refer to Figure 3 , which shows a computer device for a method of text-based clipped video according to the present invention, specifically including the following: The above computer device 12 is presented in the form of a general computing device, and the components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0066] The bus 18 represents one or more of several types of bus 18 structures, including a memory bus 18 or a memory controller, a peripheral bus 18, a graphics acceleration port, a processor, or a local bus 18 using any bus 18 structure in a variety of bus 18 structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus 18, Micro Channel Architecture (MAC) bus 18, Enhanced ISA bus 18, Video Electronics Standards Association (VESA) local bus 18, and Peripheral Component Interconnect (PCI) bus 18.

[0067] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0068] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computing device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading from and writing to a non-removable, non-volatile magnetic medium (commonly known as a "hard disk drive"). Although Figure 3 not shown in Figure 3 , a disk drive for reading from and writing to a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading from and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) can be provided. In these cases, each drive can be connected to the bus 18 by one or more data media interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42 that are configured to perform the functions of the embodiments of the present invention.

[0069] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the memory. Such program modules 42 include—but are not limited to—an operating system, one or more application programs, other program modules 42, and program data, and an implementation of a network environment may be included in each or some combination of these examples. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.

[0070] The computing device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, a camera, etc.), and can also communicate with one or more devices that enable healthcare personnel to interact with the computing device 12, and / or communicate with any device that enables the computing device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Also, the computing device 12 can communicate with one or more networks (such as a local area network (LAN)), a wide area network (WAN), and / or a public network (such as the Internet) through the network adapter 20. As shown, the network adapter 20 communicates with other modules of the computing device 12 through the bus 18. It should be understood that although Figure 3 not shown in Figure 3 , other hardware and / or software modules can be used in conjunction with the computing device 12, including but not limited to: microcode, device drivers, redundant processing unit 16, external disk drive arrays, RAID systems, tape drives, and data backup storage system 34, etc.

[0071] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing a method for text-clipped video provided by the embodiments of the present invention.

[0072] That is, when the above-mentioned processing unit 16 executes the above program, it realizes: obtaining the video to be processed and the corresponding explanatory text, extracting image frames from the video to be processed according to a preset frame number interval, and determining the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; Selecting an unprocessed target explanatory word from the explanatory text according to the natural paragraph order, and determining the text feature vector corresponding to the target explanatory word; Determining the target image frame with the highest similarity to the target explanatory word vector according to the text feature vector and the image feature vector; Performing video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video.

[0073] In the embodiments of the present invention, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it realizes a method for text-clipped video provided by all embodiments of the present application: That is, when the program is executed by a processor, it realizes: obtaining the video to be processed and the corresponding explanatory text, extracting image frames from the video to be processed according to a preset frame number interval, and determining the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; Selecting an unprocessed target explanatory word from the explanatory text according to the natural paragraph order, and determining the text feature vector corresponding to the target explanatory word; Determining the target image frame with the highest similarity to the target explanatory word vector according to the text feature vector and the image feature vector; Performing video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video.

[0074] Any combination of one or more computer-readable media may be employed. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present document, a computer-readable storage medium may be any tangible medium that contains or stores a program which can be used by or in connection with an instruction execution system, apparatus, or device.

[0075] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0076] Computer program code for carrying out operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages - such as Java, Smalltalk, C++ - and also including conventional procedural programming languages - such as the "C" programming language or similar programming languages. The program code may execute entirely on the healthcare provider's computer, partly on the healthcare provider's computer, as a stand-alone software package, partly on the healthcare provider's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the healthcare provider's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). Each embodiment in this specification is described in a progressive manner, with the emphasis of each embodiment being on the differences from other embodiments. The same or similar parts among the various embodiments may be referred to each other.

[0077] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0078] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0079] The above has introduced in detail a method and device for text-clipped videos provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for text-clipping-based videos, characterized in that, Including the steps: Obtain the video to be processed and the corresponding explanatory text, extract image frames from the video to be processed according to a preset frame number interval, and determine the image feature vectors corresponding to the image frames; wherein, the number of the extracted image frames is at least 1 frame; Select an unprocessed target explanatory word from the explanatory text according to the natural paragraph order, and determine the text feature vector corresponding to the target explanatory word; Determine the target image frame with the highest similarity to the target explanatory word vector according to the text feature vector and the image feature vector; Perform video interception on the video to be processed based on the timing position corresponding to the target image frame to generate a target clipped video.

2. The method according to claim 1, characterized in that The step of extracting image frames from the video to be processed according to a preset frame number interval and determining the image feature vectors corresponding to the image frames includes: Extract image frames from the video to be processed according to a preset frame number interval, and construct an image database; Input each image frame in the image database into a multi-modal alignment model, and calculate the image feature vectors corresponding to each image frame.

3. The method according to claim 1, characterized in that, The step of selecting an unprocessed target explanatory word from the explanatory text according to the natural paragraph order and determining the text feature vector corresponding to the target explanatory word includes: Segment the explanatory text into multiple explanatory words according to semantic integrity, and construct an explanatory word processing sequence; wherein, each explanatory word has a processing identifier, and the processing identifier includes an unprocessed identifier and a processed identifier; Select the first explanatory word with an unprocessed identifier from the processing sequence as the target explanatory word according to the original text arrangement order of the explanatory text; Input the target explanatory word into a multi-modal alignment model, and calculate the text feature vector corresponding to the target explanatory word.

4. The method according to claim 3, wherein It further includes: Mark the status identifier of the target explanatory word for which the target clipped video has been generated in the processing sequence as processed.

5. The method according to claim 1, wherein The step of determining the target image frame with the highest similarity to the target explanatory word vector according to the text feature vector and the image feature vector includes: Input the text feature vector and the image feature vector into a multi-modal alignment model, and calculate the similarity between the text feature vector of the target explanatory word and the image feature vectors of each image frame; Perform a descending order sorting on the vector similarities, and select the image frame with the highest similarity to the target explanatory word vector as the target image frame according to the sorting result.

6. The method according to claim 1, characterized in that The step of performing video interception on the video to be processed based on the corresponding timing position of the target image frame to generate a target clipped video includes: Obtain the first timing position of the target image frame in the video to be processed; Set the target image frame as the timing midpoint of the target clipped video; Intercept a video segment with a preset duration from the video to be processed according to the first timing position and the timing midpoint as the target clipped video.

7. The method according to claim 6, wherein The step of intercepting a video segment with a preset duration from the video to be processed according to the first timing position and the timing midpoint as the target clipped video includes: Calculate the start time point and the end time point of the target clipped video according to the timing midpoint and the preset duration; Calculate the starting and ending capture timing positions in the video to be processed based on the starting time point and the ending time point; Capture a video segment from the video to be processed as the target clip video based on the starting and ending capture timing positions; 8. The method according to claim 1, characterized in that, The step of performing video capture on the video to be processed based on the timing position of the target image frame to generate a video segment to be processed includes: Obtain the second timing position of the target image frame in the video to be processed; Obtain a transition frame corresponding to the target image frame from the video to be processed, and obtain the timing position corresponding to the transition frame; Capture a video segment from the video to be processed as the target clip video based on the second timing position and the timing position corresponding to the transition frame; 9. The method according to claim 8, wherein The step of capturing a video segment from the video to be processed as the target clip video based on the second timing position and the timing position of the transition frame includes: Calculate the starting transition frame and the ending transition frame with the closest timing distance to the target image frame based on the second timing position and the timing position corresponding to the transition frame respectively; Capture a video segment from the video to be processed as the target clip video based on the timing position of the starting transition frame and the timing position of the ending transition frame; 10. An apparatus for text-clipping based video, characterized in that, Includes: An image feature extraction module, configured to extract image frames from the video to be processed at preset frame intervals and determine the image feature vectors corresponding to the image frames; A text feature extraction module, configured to select an unprocessed target commentary from the commentary text according to the natural paragraph order and determine the text feature vector corresponding to the target commentary; A cross-modal matching module, configured to determine the target image frame with the highest similarity to the target commentary vector according to the text feature vector and the image feature vector; A video generation module, configured to perform video capture on the video to be processed based on the timing position corresponding to the target image frame to generate a target clip video.

Citation Information

Patent Citations

  • Live broadcast-based video pushing method and device and computer readable storage medium

    CN110198456A

  • Method and device for generating plot explanation short video and electronic equipment

    CN114222196A

  • Video editing method and device based on script, equipment and medium

    CN114245203A

  • Video editing method and device and computer readable storage medium

    CN117097944A

  • Video generation method and device, computer equipment and storage medium

    CN118540542A