A script-based multi-modal feature matching video clip method and system

By using script-based multimodal feature matching technology, automated video editing methods solve the problems of time-consuming and labor-intensive traditional short video production, enabling intelligent editing to generate high-quality videos that can adapt to different content needs.

CN117880443BActive Publication Date: 2026-07-31CHENGDU SOBEY DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU SOBEY DIGITAL TECH CO LTD
Filing Date
2023-12-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional short video production is time-consuming and labor-intensive, and existing automatic production methods based on video templates are limited by fixed templates and cannot flexibly adapt to different content needs.

Method used

By matching the text vector features of the video production script with the multimodal features of the candidate videos, and using an attention mechanism to align and fuse the features, the optimal video segment is recommended and intelligently edited, and the final video is generated by combining user preferences.

Benefits of technology

It enables the automatic matching of templates that match the video theme based on the provided video script, and intelligent editing to generate high-quality videos, reducing manpower consumption and improving production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117880443B_ABST
    Figure CN117880443B_ABST
Patent Text Reader

Abstract

This invention provides a script-based multimodal feature matching video editing method, comprising: acquiring a video production script and candidate videos; extracting text vector features from the video production script and segmenting the candidate videos and extracting multimodal video vector features from each video segment; aligning and fusing the features of the video production script and candidate videos based on an attention mechanism; matching video segments with optimal video vector features according to the text vector features of the video production script; and performing editing one by one based on the matched video segments; recommending matching video templates based on the edited video segments; adding the content of the video production script to the video templates and merging them with the edited video segments to obtain the finished video. This invention can achieve intelligent tag extraction and intelligent editing, and automatically match templates that match the video theme to synthesize a finished video, provided only with a video production script and multiple candidate videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a script-based multimodal feature matching video editing method and system. Background Technology

[0002] Traditional short video production and editing methods mainly include the following steps: 1. Video content planning and script writing; 2. Shooting footage or finding existing content as usable material; 3. Using video editing tools to edit the footage into a finished video according to the script requirements. Producing such a video takes a considerable amount of time, and each step requires significant manpower.

[0003] There is an existing method for automatically producing short videos based on video templates. This involves pre-editing a set of fixed video clips for a specific scene or type. Users then simply replace the video and audio clips in this set with their own prepared materials and input pre-prepared text content to complete the short video production. While this method reduces manpower and time consumption, the format of the produced video is limited by the selected video template, including the video playback area, video size, and transition animations. Essentially, the video script is fixed, and the selection of materials and application scenarios are also limited. For example, templates related to the Mid-Autumn Festival can only be used to create videos related to Mid-Autumn Festival activities. If other content is to be produced, additional templates need to be added, thus increasing the workload of the template creators. Summary of the Invention

[0004] To address the problems existing in the prior art, a script-based multimodal feature matching video editing method and system is provided. This method and system achieves intelligent editing by matching the text vector features of the video production script file with videos in the media library.

[0005] The first aspect of this invention proposes a script-based multimodal feature matching video editing method, comprising:

[0006] Obtain the video production script and candidate videos;

[0007] Extract text vector features from video production scripts and segment candidate videos to extract multimodal video vector features from each video segment;

[0008] Based on the attention mechanism, the features of the video production script and candidate videos are aligned and fused. According to the text vector features of the video production script, the video segment with the optimal video vector features is matched, and the editing is completed one by one according to the matched video segment.

[0009] Recommend matching video templates based on the edited video clips;

[0010] The video production script is added to the video template and then combined with the edited video clips to obtain the finished video.

[0011] Furthermore, the segmentation method includes: dividing the segments based on their duration, resolution, or vector features.

[0012] Furthermore, based on cue learning and domain adaptation fine-tuning, the multimodal pre-trained model is trained. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features from each video segment. The multimodal video vector features include text, images, and sound.

[0013] Furthermore, it also includes: evaluating the quality of video segments based on semantic information of video content, and extracting high-quality segments that represent the main content of the video by adding the user's personalized preferences and aggregating them to obtain a video summary.

[0014] Furthermore, the method for recommending video templates is as follows: based on the theme of the edited video clips, templates from the template library that have a high degree of matching with the features or tags of the edited video clips are recommended.

[0015] Furthermore, during the video segment synthesis process, packaging materials are determined based on the video template and incorporated into the video. These packaging materials include subtitles, establishing shots, transitions, special effects, and / or textures.

[0016] A second aspect of the present invention provides a script-based multimodal feature matching video editing system, comprising:

[0017] The script input module is used to obtain the video production script provided by the user;

[0018] The video production module is used to extract multimodal features from candidate videos based on the text vector features extracted from the video production script. It then combines video summarization extraction with user-personalized preferences and intelligent editing to match the intelligently edited video segments with recommended video templates and synthesize them into finished videos according to the content order of the video production script.

[0019] Furthermore, the video production module includes a scene segmentation and merging module and a multimodal embedding module; wherein,

[0020] The scene segmentation and merging module is used to divide a continuous video stream into independent video segments based on scene transitions.

[0021] The multimodal embedding module trains the multimodal pre-trained model based on cue learning and domain adaptation fine-tuning. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features from each video segment.

[0022] Furthermore,

[0023] The video production module includes a cross-modal feature fusion module and a score prediction module, wherein,

[0024] The cross-modal feature fusion module uses an attention mechanism to achieve cross-modal feature fusion to obtain vector features with better representation, enabling feature fusion and alignment between video production scripts and candidate videos;

[0025] The score prediction module matches video segments with the optimal video vector features based on the text vector features of the video production script.

[0026] Furthermore, the video production module also includes a video synthesis module, which completes the editing based on the matched video segments, and obtains the matching template based on the cross-modal recall and ranking algorithm using multimodal features. The video production script content is added to the video template and combined with the editing results to synthesize the finished video according to the order of the video production script content.

[0027] Compared with the existing technology, the beneficial effects of adopting the above technical solution are as follows: the present invention can achieve intelligent tag extraction and intelligent editing, and automatically match templates that match the video theme to synthesize a finished video, provided that only a video production script (lyrics) and multiple candidate videos are provided. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the script-based multimodal feature matching video editing method proposed in this invention.

[0029] Figure 2 This is a diagram of the script-based multimodal feature matching video editing framework proposed in this invention.

[0030] Figure 3 This is a schematic diagram of a specific scenario of the present invention.

[0031] Figure 4 This is a flowchart illustrating the process of obtaining video template recommendation tags in one embodiment of the present invention. Detailed Implementation

[0032] The embodiments of this application are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar modules or modules having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. Rather, the embodiments of this application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0033] To address the shortcomings of traditional automated video production methods, this invention proposes a script-based multimodal feature matching video editing method. Given a user-provided script text (lyrics) and multiple candidate videos, the method achieves intelligent tag extraction and intelligent editing, automatically matching templates that match the video theme to synthesize a finished video. The specific solution is as follows:

[0034] Please refer to Figure 1 A script-based multimodal feature matching video editing method includes:

[0035] Step 1: Obtain the video production script and candidate videos; establish a multimodal semantic feature extraction and representation framework, extract the text vector features of the video production script, and segment the candidate videos and extract the multimodal video vector features of each video segment.

[0036] The video editing method proposed in this embodiment relies on a user-provided video production script and candidate videos. The video production script typically includes information such as camera movement, shot size, duration, and scene description.

[0037] In practical applications, long and short video datasets that match business scenarios can be pre-built to collect editing materials, including images and video footage.

[0038] In this embodiment, a schema-guided multimedia content semantic feature extraction and representation method is adopted. This method is applicable to different domains and modalities, fundamentally improving the generalization of the algorithm model to adapt to more business scenarios. In the context of video semantic understanding, which inherently possesses multi-domain (news, sports, etc.) and multi-modal (text, image, sound, etc.) characteristics, research on a unified representation framework will also contribute to the subsequent updates and iterations of the entire algorithm model.

[0039] To better facilitate video editing, candidate videos are first segmented before feature extraction. In one embodiment, segmentation is performed based on segment duration, segment resolution, or segment vector features. After video segmentation, multimodal video vector feature extraction is executed.

[0040] like Figure 2 The diagram shows the script-based multimodal feature matching video editing framework proposed in this embodiment. The scene segmentation and merging module and the multimodal embedding module mainly implement the extraction and representation of semantic features of multimedia content. Specifically,

[0041] The scene segmentation and merging module is used to segment candidate videos into independent video segments based on scene transitions from a continuous video stream. The multimodal embedding module trains a multimodal pre-trained model using techniques such as cue learning and domain adaptation fine-tuning. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features (including vector features from image, audio, and other modalities) from each video segment.

[0042] Step 2: Align and merge the features of the video production script and candidate videos. Based on the text vector features of the video production script, match the video segments with the optimal video vector features, and complete the editing one by one based on the matched video segments.

[0043] exist Figure 2 The diagram illustrates a script-based multimodal feature matching video editing framework. This step is primarily implemented in the cross-modal feature fusion module and the score prediction module. Specifically, the cross-modal feature fusion module uses an attention mechanism to fuse cross-modal features to obtain vector features with better representation, achieving the effect of aligning and fusing features from the video production script and candidate videos. The score prediction module, based on the text vector features of the video, matches the video segment with the optimal video vector features to derive a content semantic score based on this framework, and completes the video editing based on the best matched video segment.

[0044] During this process, video summaries can be generated as needed, and the quality of extracted video segments can be reasonably evaluated based on the semantic information of the video content. Furthermore, considering user preferences, high-quality segments that represent the main content of the video are extracted and aggregated to obtain the video summary. By incorporating user preferences to constrain the video summary, error correction can be achieved.

[0045] In this embodiment, reference Figure 2 The generation of video summaries is mainly implemented in the score correction module. It supports the use of various user-defined preference modules and related algorithms to correct video segment scores based on the initial video summary, thus updating the video summary to match the user's preferences. The user-defined preference module allows for personalized settings by the user, and in one embodiment, it includes query keywords, personal preferences, fuzzy constraints, occlusion constraints, shot preferences, and others.

[0046] Step 3: Recommend matching video templates based on the edited video clips.

[0047] In this embodiment, based on the theme of the video clips generated by the editing, templates from the template library that have a high degree of matching with the features or tags of the video clips generated by the intelligent editing are recommended.

[0048] Please refer to Figure 2In this embodiment, a matching video template is recommended based on the original video material. The selection of video templates is achieved by using cross-modal recall and sorting. The features or tags contained in the text content, audio content and visual content corresponding to the video are fully considered to obtain the most appropriate template.

[0049] Step 4: Video synthesis.

[0050] In this embodiment, the video production script content is added to the video template; the transitions and special effects between the various segments of the script are applied according to the recommended template to determine the packaging materials for the intelligently recommended video, including but not limited to subtitles, empty shots, transitions, special effects, and textures. The script video is then synthesized to obtain the finished video and video summary.

[0051] like Figure 2 As shown, video summarization and template recommendation based on multimodal semantic features is a complex implementation process spanning multiple tracks. In this embodiment, it is divided into several sub-tasks for research and integrated into a unified framework, including: constructing long and short video datasets that conform to business scenarios; guiding and completing the construction of a multimodal semantic feature extraction and representation framework; implementing feature extraction and representation algorithms for different modalities (which may be expressed as latent space or entity labels); implementing video summarization algorithms based on extracted multimodal features under an unsupervised model architecture; based on the initial video summary, implementing various user-defined preference module-related algorithms to correct video segment scores and update to obtain video summaries that conform to user preferences; and implementing cross-modal recall and ranking algorithms based on extracted multimodal features to obtain templates that match the video topic.

[0052] by Figure 3 As shown, taking the "Large Conference" scenario as an example for recommending video clips in a specific scenario script, the following process is executed for video input in this scenario:

[0053] Speech Recognition: Converts speech signals into text information. Shot Segmentation: Segments a continuous video stream into independent shot segments based on scene transitions. Face Recognition: Identifies and labels different faces in a video. Meeting Scene Motion: Detects the movement of objects or people in a meeting scene. Meeting Scene Shot Size: Identifies changes in shot size within a meeting scene. Meeting Scene Angle Recognition: Identifies the camera's shooting angle in a meeting scene. Exterior Scene Motion: Detects the movement of objects or people in an exterior scene. Exterior Scene Shot Size: Identifies changes in shot size within an exterior scene. Exterior Scene Angle Recognition: Identifies the camera's shooting angle in an exterior scene. Specific Target Recognition: Identifies and tracks specific targets or objects in a video. Specific Scene Recognition: Identifies specific scenes or environments in a video. Scene Recognition: Identifies different scenes and backgrounds in a video. Human Behavior Recognition: Identifies and classifies the behavior and actions of people in a video. Audio Analysis: Analyzes and processes audio signals from a video.

[0054] Then, combined with script input, cross-modal matching of text and images is performed: cross-modal information matching and association are conducted between images, text, and audio to achieve more accurate retrieval. Finally, fusion analysis and video summarization are performed: fusion analysis and reasoning can integrate the above capabilities to achieve more advanced video summarization functions.

[0055] The script designed in this example covers 17 editable fields, including storyboard content description, voice-over text, additional conditions, scene, meeting-shot type, meeting-shot angle, meeting-shot movement, meeting-synchronous type, exterior-shot type, exterior-shot angle, exterior-shot movement, exterior-specific target, exterior-specific scene, exterior-synchronous type, on-screen characters, character behavior, and synchronous sound. The script also includes the editable range, constraints, and input and output interface design for each field.

[0056] Based on the specified large-scale conference script, an overall large-scale conference editing process was designed, which includes 15 basic modules: shot segmentation, face recognition, motion, shot size and angle recognition of conference scenes, motion, shot size and angle recognition of outdoor scenes, specific target recognition, specific scene recognition, image scene recognition, human behavior recognition, audio analysis, cross-modal matching of text and images, and fusion analysis and reasoning.

[0057] Based on existing template types and considering factors such as current technological capabilities, timeliness, resource costs, and tag credibility, a technical approach was determined: using OCR and ASR to obtain text information, and text classification, event extraction, and regular expression operations to extract tags. The specific process is as follows: Figure 4 As shown, the following content tags were ultimately identified as matching the template tags: law, economy, culture, education, sports (comprehensive event extraction), science, technology, medicine and health, entertainment and leisure, personnel appointments and removals, organizational behavior (e.g., visits, tours, research), accidents, natural disasters - earthquake disasters, natural disasters - flood disasters, natural disasters - geological disasters, natural disasters - meteorological disasters, legal events (e.g., lawsuits, interviews), festivals, and solar terms.

[0058] In one embodiment, the present invention also proposes a script-based multimodal feature matching video editing system, comprising:

[0059] The script input module is used to obtain the video production script given by the user. The video production script contains multiple video production requirements, and each video production supports information such as camera movement, shot size, duration, and scene description.

[0060] The video production module is used to extract multimodal features from candidate videos based on the text vector features extracted from the video production script. It then combines video summarization extraction with user-personalized preferences and intelligent editing to match the intelligently edited video segments with recommended video templates and synthesize them into finished videos according to the content order of the video production script.

[0061] Specifically, the video production module includes: semantic feature extraction of multimedia content based on schema guidance, a semantic feature extraction method applicable to different fields and modalities; video summarization and intelligent editing based on multimodal features and user personalized preferences, the main purpose of which is to extract high-quality segments that can represent the main content of the video based on the semantic information of the video content and the user's personalized preference settings; video template recommendation based on optimal feature matching, which selects appropriate templates for video segments generated by intelligent editing; and intelligent video synthesis, which synthesizes the edited video into a video according to script rules and template rules.

[0062] Please refer to Figure 2 In this embodiment, the video production module includes a scene segmentation and merging module and a multimodal embedding module; wherein,

[0063] The scene segmentation and merging module is used to divide a continuous video stream into independent video segments based on scene transitions. The segmentation logic includes, but is not limited to, the duration of the video segment, the resolution of the video segment, or the vector features of the video segment.

[0064] The multimodal embedding module trains the multimodal pre-trained model based on cue learning and domain adaptation fine-tuning. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features from each video segment.

[0065] In this embodiment, the video production module includes a cross-modal feature fusion module and a score prediction module, wherein,

[0066] The cross-modal feature fusion module uses an attention mechanism to achieve cross-modal feature fusion to obtain vector features with better representation, enabling feature fusion and alignment between video production scripts and candidate videos;

[0067] The score prediction module matches video segments with the optimal video vector features based on the text vector features of the video production script.

[0068] In this embodiment, the video production module also includes a video synthesis module, which completes the editing based on the matched video segments, and obtains the matching template based on the cross-modal recall and sorting algorithm using multimodal features. The video production script content is added to the video template and combined with the editing results to synthesize the finished video according to the order of the video production script content.

[0069] It should be noted that in this embodiment, the quality of the extracted video segments is also reasonably evaluated based on the semantic information of the video content; in addition, considering the user's personalized preferences, high-quality segments that can represent the main content of the video are extracted and aggregated to obtain a video summary.

[0070] Furthermore, in this embodiment, matching video templates are recommended based on the original video footage. A cross-modal recall and ranking algorithm is implemented based on extracted multimodal features to obtain templates that match the video theme, resulting in a template recommendation result that fits the video. The obtained video modules can then have transitions and effects added directly between script segments during video synthesis to generate a high-quality finished video.

[0071] In particular, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts.

[0072] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0074] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0075] In another aspect, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the script-based multimodal feature matching video editing method and system described in the above embodiments.

[0076] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the script-based multimodal feature matching video editing method and system described in the above embodiments.

[0077] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0078] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0079] For those skilled in the art, the specific meanings of the above terms in this invention can be understood according to the specific circumstances; the accompanying drawings in the embodiments are used to clearly and completely describe the technical solutions in the embodiments of this invention. Obviously, the described embodiments are some embodiments of this invention, but not all embodiments. Generally, the components of the embodiments of this invention described and shown in the accompanying drawings can be arranged and designed in various different configurations.

[0080] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A script-based multimodal feature matching video editing method, characterized in that, include: Obtain the video production script and candidate videos; Extract text vector features from video production scripts and segment candidate videos to extract multimodal video vector features from each video segment; Based on the attention mechanism, the features of the video production script and candidate videos are aligned and fused. According to the text vector features of the video production script, the video segment with the optimal video vector features is matched, and the editing is completed one by one according to the matched video segment. Recommend matching video templates based on the edited video clips; The video production script is added to the video template and combined with the edited video clips to obtain the finished video. Based on prompting learning and domain adaptation fine-tuning, the multimodal pre-trained model is trained. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features from each video segment. The multimodal video vector features include text, images, and sound. It also includes: evaluating the quality of video segments based on semantic information of video content, and extracting high-quality segments that represent the main content of the video by adding the user's personalized preferences and aggregating them to obtain a video summary.

2. The script-based multimodal feature matching video editing method according to claim 1, characterized in that, The segmentation method includes: dividing the segments based on their duration, resolution, or vector features.

3. The script-based multimodal feature matching video editing method according to claim 1, characterized in that, The method for recommending video templates is as follows: based on the theme of the edited video clips, recommend video templates from the template library that have a high degree of matching with the features or tags of the edited video clips.

4. The script-based multimodal feature matching video editing method according to claim 1, characterized in that, During the video segment synthesis process, packaging materials are determined based on the video template and then incorporated into the video. These packaging materials include subtitles, empty shots, transitions, special effects, and / or textures.

5. A script-based multimodal feature matching video editing system, characterized in that, include: The script input module is used to obtain the video production script provided by the user; The video production module is used to extract multimodal features from candidate videos based on the text vector features extracted from the video production script, and combine them with video summary extraction and intelligent editing based on user personalized preferences. The intelligently edited video segments are matched with recommended video templates and synthesized into finished videos according to the content order of the video production script. The video production module includes a scene segmentation and merging module and a multimodal embedding module; wherein... The scene segmentation and merging module is used to divide a continuous video stream into independent video segments based on scene transitions; The multimodal embedding module trains the multimodal pre-trained model based on prompting learning and domain adaptation fine-tuning. The trained multimodal model is then used to extract text vector features from the video production script and multimodal video vector features from each video segment. The video production module includes a cross-modal feature fusion module and a score prediction module, wherein, The cross-modal feature fusion module uses an attention mechanism to achieve cross-modal feature fusion to obtain vector features with better representation, enabling feature fusion and alignment between video production scripts and candidate videos; The score prediction module matches video segments with the optimal video vector features based on the text vector features of the video production script.

6. The script-based multimodal feature matching video editing system according to claim 5, characterized in that, The video production module also includes a video synthesis module, which completes the editing based on the matched video segments, and obtains the matching template based on the cross-modal recall and ranking algorithm using multimodal features. The video production script content is added to the video template and combined with the editing results to synthesize the finished video according to the order of the video production script content.