Video editing method and system based on multi-modal large model collaboration
Through the video editing method of multimodal large model collaboration, the problems of inaccurate product segmentation and inaccurate matching of highlights and materials in single-modal editing methods are solved, and efficient and accurate video editing and output are achieved, which improves editing efficiency and audience interaction effects.
Patent Information
- Application Number
- CN202511287057.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-10
AI Technical Summary
Existing automated video editing methods rely on a single modality, resulting in inaccurate product segmentation, out-of-sync highlight extraction and image processing, and inaccurate material matching.
A video editing method that collaborates with multimodal large models is adopted. Video is segmented through audio feature extraction and semantic integrity judgment. Scenes are separated and verified by combining language and visual large models. Multi-dimensional highlight sentences are extracted and weighted, product visual close-ups are located and inserted, and materials are intelligently matched and integrated for output.
It improves the accuracy of product scene segmentation, enhances the audience appeal of highlight clips, improves the hit rate of material matching, ensures that the video output format matches the platform characteristics, and greatly improves editing efficiency.
Smart Images

Figure CN120786154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a video editing method and system based on multimodal large model collaboration. Background Art
[0002] With the rapid development of live e-commerce, massive amounts of live videos need to be quickly edited into highlight clips to meet the needs of secondary dissemination. Traditional manual editing suffers from low efficiency, high costs, and inconsistent standards. The emergence of automated editing effectively alleviates these problems. Existing automated editing methods often rely on a single modality (such as audio or visual), resulting in inaccurate product segmentation, out-of-sync highlight extraction and image quality, and imprecise material matching. Summary of the Invention
[0003] Purpose of the invention: The purpose of the present invention is to provide a video editing method and system based on the collaboration of multimodal large models; it can solve the problems of existing automated editing methods relying on a single modality, resulting in inaccurate product segmentation, highlight extraction and picture synchronization, and inaccurate material matching.
[0004] Technical solution: To solve the above technical problems, according to one aspect of the present invention, more specifically, a video editing method and system based on multimodal large model collaboration, specifically comprising the following steps: S1. Video preprocessing and intelligent segmentation: Preprocess the original video, cut it into small segments, ensure that each segment ends with a complete sentence, output the timestamp and audio text segment of each segment, and store them in structured data; S2. Multimodal scene segmentation and verification: Each block, i.e., audio text fragment and its timestamp, is input into the language model. The scene is initially segmented according to the rules and then a sprite image is constructed. This is then input into the visual model to verify whether the product has switched. Finally, the scene time, title, and description are output. S3. Product highlight sentence extraction and ranking: The scene text is input into the language model, and the multi-dimensional highlight information is combined to extract the highlight sentences. Audience interaction data and product type preset templates are introduced to adjust the sentence weights, and the sentence is sorted by weight to form a highlight clip timeline. S4. Product visual close-up positioning and insertion: The scene video is divided, labeled, and numbered. The visual model is input to obtain the grid number of the product to locate the product. The close-up parameters are dynamically adjusted according to the product type. The close-up is inserted into the corresponding interval of the highlight sequence through fade-in and fade-out transitions. S5. Intelligent Keyword Extraction and Material Matching: Keywords are extracted and sorted by frequency and importance. Text descriptions are generated for each material, including details, purpose, and emotion. Keywords and material descriptions are converted into vectors. By matching materials with the highest similarity, dynamic materials with high-frequency occurrences of preset keywords are prioritized and inserted at the corresponding time points. S6. Video integration packaging and adaptive output: Integrate the content of the main track, close-up track, and material track, adjust the style and duration according to the output platform, detect and adjust the screen occlusion before output, and finally output the video and metadata.
[0005] Furthermore, the step S1 specifically includes the following steps: S11. Audio feature extraction: Perform audio separation on the original video to extract the sound wave intensity sequence and speech pause features; S12, Semantic Integrity Assessment: Call the speech recognition model to convert the audio into text, analyze the sentence boundaries of the text using the language model, and detect the time points corresponding to punctuation marks, including periods and exclamation marks; S13. Block rule execution: Based on the sound wave intensity, low-intensity segments are prioritized as breakpoints and sentence end times. The video is cut into small segments within a preset time range, ensuring that each segment ends with a complete sentence. S14. Block result storage: Output the start and end timestamps of each block and the corresponding audio text segment, and store them as structured data.
[0006] Furthermore, the step S2 specifically includes the following steps: S21. Input each segmented text into the language model, and perform scene segmentation based on preset rules. The following strategies are used for the segmentation results: if no product is segmented or only one product is segmented, it is considered that the product explanation is not completed, and the next subtitle data will be spliced in. If two products are segmented and the second product explanation does not exceed two minutes, it is considered that it has not yet ended. A delay of 2 minutes is added to ensure segmentation stability. If two products are segmented and the second product explanation exceeds two minutes, the first product explanation is considered to be completed, and the first segmentation result is output. The subtitles belonging to the second product are spliced in. If multiple segmentation results appear, only the last two results are retained, and the rest are output. If multiple short segmentation results of less than 2 minutes appear, only the last two results are retained, and the rest are merged into one result and named "complex scene"; S22, extracting frame images from the video at a predetermined time before and after the segmentation point of the preliminary segmentation result at a frequency of one frame per second, and splicing them into sprite images arranged horizontally with a frame spacing of 10 pixels in chronological order; S23. Input the sprite image into the visual model to identify product switching features. If the product switching is visually confirmed, retain the initial classification result. S24: If the segmentation has not been switched, extend the segmentation point backward according to the continuity of the slow image, and re-execute steps S21-S23 until the segmentation is visually confirmed to be accurate; S25. Output the scene text of each scene including the start and end time, scene title, and content introduction.
[0007] Further, the step S3 specifically comprises the following steps: S31, inputting the scene text into a language large model, extracting high-light sentences in combination with a preset high-light dimension, and outputting the start time, end time and content of the sentences; S32, introducing auxiliary data to optimize high-light sorting, including combining the keywords of live broadcast at the time of the barrage, the peak time of likes, adding the audience interaction weight of the weight of the sentences mentioning the corresponding content, and the commodity type weight of the preset sorting template according to the commodity type; S33, sorting the high-light sentences according to the weight calculation results in the order of the template to form a high-light segment basic timeline.
[0008] Further, when the step S4 inserts the commodity close-up positioning, the video frame is divided into a 50x50 pixel grid picture, the center of each grid is marked with a red dot with a black outline and numbered, a marked frame image is formed, the marked frame image is input into a visual large model, the grid number where the commodity is located is output in combination with the high-light sentence content, the close-up parameters are dynamically adjusted according to the commodity type, the generated commodity close-up picture is inserted into the time interval of the corresponding high-light sentence, and the picture is smoothly transitioned through fade-in and fade-out transition.
[0009] Further, when the step S5 extracts intelligent keywords and matches materials, a language large model and precision prompt words are combined to extract keywords: input the scene text and high-light sentences, the prompt words are set as "extract commodity name, core buying point, preferential information, emotional tendency", the extracted keywords are sorted according to the appearance frequency and preset importance, the materials are preprocessed to generate description texts according to "detail description + purpose + emotional label", the keywords and material text descriptions are converted into vectors, the matching degree is calculated through cosine similarity, the top three materials are selected according to the matching degree, if the appearance frequency of the preset keywords is greater than the preset frequency, the dynamic stickers of the countdown animation are preferentially matched, and the matched materials are inserted into the corresponding positions according to the timeline.
[0010] Further, in the step S6, when the video integration packaging and adaptive output are performed, the close-up track: commodity close-up picture and the main track: high-light sentence corresponding video segmentation picture are superimposed, and then the material track: matched material is inserted according to the timestamp, according to the preset style parameters and time length parameters of the output platform, the picture is detected by the visual large model before output, if there is an occlusion, the material position is automatically adjusted, and the final video and metadata are output.
[0011] According to another aspect of the present invention, a video editing system based on multimodal large model collaboration is provided, characterized in that: the system is used to implement the above-mentioned video editing method based on multimodal large model collaboration, and includes: a video preprocessing and intelligent segmentation module, a multimodal scene separation and verification module, a product highlight sentence extraction and sorting module, a product visual close-up positioning and insertion module, an intelligent keyword extraction and material matching module, and a video integration packaging and adaptive output module; Video preprocessing and intelligent segmentation module: This module preprocesses the original video, cuts it into small chunks, ensures that each chunk ends with a complete sentence, outputs the timestamp and audio text fragment of each chunk, and stores them in structured data. Multimodal scene segmentation and verification module: This module inputs each segment into the language model, initially divides the scene according to the rules, and then constructs the sprite map. This module inputs the visual model to verify whether the product has been switched, and finally outputs the scene time, title, and description. Product highlight sentence extraction and sorting module: This module is used to input scene text into the language model, extract highlight sentences based on multi-dimensional highlight information, introduce audience interaction data and product type preset templates to adjust sentence weights, and sort by weight to form a highlight clip timeline; Product visual close-up positioning and insertion module: This module is used to divide, mark, and number scene videos, input the visual model to obtain the grid number where the product is located to locate the product, dynamically adjust the close-up parameters based on the product type, and insert the close-up into the corresponding interval of the highlight sequence through fade-in and fade-out transitions; Intelligent keyword extraction and material matching module: This module extracts keywords and sorts them by frequency and importance. It generates a text description of each material, including details, purpose, and emotion. It converts keywords and material descriptions into vectors. By matching materials with the highest similarity, it prioritizes dynamic materials with high-frequency occurrences of preset keywords and inserts them into the corresponding time points. Video integration packaging and adaptive output module: used to integrate the content of the main track, close-up track, and material track, adjust the style and duration according to the output platform, detect and adjust the screen occlusion before output, and finally output the video and metadata.
[0012] Beneficial effects: 1. Through a combined strategy of "audio feature extraction + semantic integrity assessment + quantitative chunking rules," this approach addresses the issues of sentence truncation and timestamp confusion that commonly occur in traditional video chunking. Specifically, by combining sound wave intensity, speech pause characteristics, and sentence boundaries identified by the language giant model, each chunk is ensured to end with a complete sentence. Chunk duration and structured storage are standardized, providing unified and standardized input for the subsequent collaborative processing of the language giant model and the visual giant model, avoiding processing errors caused by fragmented or incomplete input.
[0013] 2. Through the multi-modal collaborative mechanism of "language large model initial division + visual wizard graph verification + dynamic adjustment cycle", the problem of inaccurate product segmentation caused by relying on only audio or visual single mode in existing methods is broken. After the language large model divides the scene based on product name, explanation duration and other rules, the visual large model verifies whether the product actually switches through the time-dimension continuous wizard graph. If it does not switch, it adjusts through the cycle of "extending the segmentation point → re-dividing → verifying again", until the visual confirms the accurate segmentation. This process effectively avoids the mis-segmentation situation of "the anchor mentions the next product in advance, but the screen does not switch" or "the screen switches, but the audio is not synchronized", improves the product scene segmentation accuracy, and ensures the complete explanation process of each scene corresponding to a single product.
[0014] 3. Through the mechanism of "multi-dimensional highlight extraction + weight optimization sorting", the problem of traditional highlight extraction relying only on text keywords and being out of touch with audience interest is solved. On the one hand, the language large model extracts highlight sentences based on pre-set dimensions such as product core selling points and discount information; on the other hand, it introduces audience interaction data and product type templates to adjust the weight of sentences, so that the highlight sequence contains not only product key information, but also audience focus and cognitive logic.
[0015] 4. Through the "grid marking positioning + dynamic parameter adaptation" technology, the problem of inaccurate product location description and incoordination between close-up and picture in the visual large model is solved. The video frame is divided into 50x50 pixel grids and marked with numbered red dots. The visual large model only needs to output the serial number to accurately locate the product, avoiding "illusion phenomenon"; at the same time, the close-up parameter is dynamically adjusted according to the product type, and the fade-in and fade-out transition is realized to achieve smooth transition between close-up and main picture. This process ensures that the product close-up highlights the details without damaging the integrity of the picture, so that the audience can clearly capture the key features of the product, and the visual expression is comparable to manual fine editing, but the efficiency is greatly improved.
[0016] 5. Through the mechanism of "language large model keyword extraction + material text description vector matching", the problem of traditional TF-IDF algorithm keyword extraction limitations and material matching interference is overcome. After the language large model extracts accurate keywords such as product name and core selling points based on fine-tuning prompt words, the material is converted into a vector through text processing of "detail description + use + emotional label", and then matched through cosine similarity, reducing the interference data of direct image vector matching, and improving the material matching hit rate by more than 40%. At the same time, when preset keywords appear frequently, the preset stickers are matched first, so that the packaging material is highly consistent with the content focus of the scene, enhancing the information transmission efficiency and visual appeal of the video.
[0017] 6. Through the "multi-track integration + platform adaptive adjustment + output pre-check" mechanism, the problem of single output format and mismatch with platform characteristics of existing editing methods is solved. Multi-track integration ensures the hierarchical cooperation of high light segments, product close-ups, and packaging materials, avoiding picture confusion; adjusting style parameters according to the output platform makes the video conform to the platform transmission characteristics; before output, visual large model detects picture occlusion and automatically adjusts to ensure the integrity and professionalism of the final video. This process improves the secondary dissemination conversion rate of the video on different platforms by more than 30%, and realizes rapid output within 10 minutes after live broadcast, greatly improving the editing efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a method flow diagram. DETAILED DESCRIPTION
[0019] In order to make the technical scheme of the present application clearer, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0020] EMBODIMENT Video editing method based on multi-modal large model cooperation.
[0021] Step 1: Video preprocessing and intelligent blocking: Preprocess the original video, cut the video into small blocks, ensure that the end of each small block is a complete sentence, output the timestamps of each block, audio text segments and store them as structured data. Specifically, the following steps are included: 1. Audio feature extraction: Separate the audio from the original video, extract the sound intensity sequence and speech pause features, such as the position of the silent segment length ≥ 0.5 seconds.
[0022] 2. Semantic integrity judgment: Call the speech recognition model such as Volcano ASRapi to convert the audio to text, analyze the sentence boundaries of the text through the language large model, and detect the time points corresponding to the punctuation including period, exclamation point, etc.
[0023] 3. Block rule execution: Combine the sound intensity, low intensity segment is preferred as the breakpoint and sentence end time, cut the video into small blocks with a preset time range, ensure that the end of each small block is a complete sentence, such as "This lipstick color will be introduced here".
[0024] 4. Block result storage: Output the start and end timestamps of each block, the corresponding audio text segment, and store it as structured data such as JSON format.
[0025] By the combination strategy of "audio feature extraction + semantic integrity judgment + quantitative block rule", the problems of sentence truncation and timestamp disorder in traditional video block are solved. Specifically, by combining the sound intensity, speech pause features and language model recognition of sentence boundaries, it ensures that each block ends with a complete sentence, and the block duration and structured storage standardization provide a unified and standardized input for subsequent language models and visual models, avoiding processing errors caused by fragmented or incomplete input.
[0026] Second step, multi-modal scene separation and verification: input each block, i.e. audio text segment and its timestamp, into the language model, and reconstruct the sprite according to the rules, input the visual model to verify whether the goods are switched, and finally output the scene time, title and synopsis: 1. Input each block text of the block into the language model, such as LLaMA series, and divide the scene based on the preset rules, such as keywords like product name and function description. For the segmentation results, the following strategies are adopted: if a product or only one product is segmented, it is considered that the product explanation has not ended, and the next subtitle data will be spliced; if two products are segmented and the second product explanation does not exceed two minutes, it is considered that it has not ended, and delay for two minutes to ensure stable segmentation; if two products are segmented and the second product explanation exceeds two minutes, it is considered that the first product explanation has ended, and the first segmentation result is output, the second product's subtitle is spliced continuously, if multiple segmentation results appear, only the last two results are retained, the rest are output, if multiple continuous short segmentation (less than 2 minutes) results appear, only the last two results are retained, the rest are merged into one result and named as "complex scene".
[0027] 2. Extract frame images from the video 30 seconds before and after the preliminary segmentation result at a frequency of one frame per second, splice them into a horizontal arrangement in chronological order, and convert the time dimension into the spatial dimension.
[0028] 3. Input the sprite into the visual model, such as the improved CLIP model, to identify the product switching features, such as the main product changing from "lipstick" to "foundation", if the visual confirms the product switching, the preliminary result is retained.
[0029] 4. If not, extend the segmentation point backward by 30-60 seconds according to the continuity of the slow motion, and re-execute steps 1-3 until the visual confirms the accurate segmentation.
[0030] 5. Output the scene text including the start and end time, scene title and content synopsis.
[0031] Through the multi-modal collaborative mechanism of "language large model initial segmentation + visual sprite verification + dynamic adjustment cycle", the problem of inaccurate product segmentation caused by relying on only audio or visual single mode in existing methods is broken through. After the language large model initially segments the scene based on product name, explanation duration, etc., the visual large model verifies whether the product actually switches through the time-dimensionally continuous sprite. If it does not switch, it adjusts through the cycle of "extending the segmentation point → re-segmenting → verifying again", until the visual confirms the accurate segmentation. This process effectively avoids the mis-segmentation of "the anchor mentions the next product in advance, but the screen does not switch" or "the screen switches, but the audio is not synchronized", improves the accuracy of product scene segmentation, and ensures that each scene corresponds to the complete explanation process of a single product.
[0032] Third step, product highlight sentence extraction and sorting: input the scene text into the language large model, extract the highlight sentence combined with multi-dimensional highlight information, introduce audience interaction data and product type preset template to adjust the weight of the sentence, and sort according to the weight to form the highlight segment timeline. Specifically, the following steps are included: 1. Input the scene text into the language large model, extract the highlight sentence combined with the preset highlight dimensions such as product core selling points, discount information, and user concerned question answers, output the start and end time of the sentence and the content, such as "buy 2 get 1 free for this cream, only today", timestamp 00:05:20-00:05:30.
[0033] 2. Introduce auxiliary data to optimize highlight sorting, including combining the live chat keywords such as "price" and "effect", the peak time of likes, increasing the weight of the audience interaction weight of the sentence mentioning the corresponding content, and sorting according to the preset sorting template of the product type, such as clothing: style → material → matching → discount; electronic product: function → performance → price → after-sales service) product type weight.
[0034] 3. According to the weight calculation result, sort the highlight sentences according to the template order to form the basic timeline of the highlight segment.
[0035] Through the "multi-dimensional highlight extraction + weight optimization sorting" mechanism, the problem of traditional highlight extraction relying only on text keywords and being out of touch with audience interest is solved. On the one hand, the language large model extracts highlight sentences combined with product core selling points, discount information, etc. preset dimensions; on the other hand, it introduces audience interaction data and product type template to adjust the weight of the sentence, so that the highlight sequence contains not only product key information, but also audience attention focus and cognitive logic.
[0036] In the fourth step, the visual close-up of the commodity is positioned and inserted. The scene video is divided, labeled and numbered, the visual large model is inputted to obtain the grid number where the commodity is located, the close-up parameters are dynamically adjusted according to the type of the commodity, and the close-up is inserted into the corresponding interval of the highlight sequence through fade-in and fade-out transition. Specifically, the video frame is divided into a 50x50 pixel grid picture, the center of each grid is marked with a black-dashed red dot (diameter 5 pixels) and numbered, forming a labeled frame image, the labeled frame image is inputted into the visual large model, combined with the highlight sentence content, such as "look at the texture of the zipper", the grid number where the commodity is located is outputted, such as "number 23", the close-up parameters are dynamically adjusted according to the type of the commodity: small commodities (such as jewelry, lipstick): enlarged to 60%-70% of the screen ratio, 2-3 seconds long, highlighting details; large commodities (such as furniture, home appliances): enlarged to 40%-50% of the screen ratio, 3-5 seconds long, taking into account the whole and the part, the generated close-up picture of the commodity is inserted into the time interval of the corresponding highlight sentence, such as the sentence "zipper texture" corresponds to 00:08:10-00:08:15, the close-up is inserted at 00:08:12-00:08:14, and the fade-in and fade-out transition is realized to achieve smooth transition of the picture.
[0037] Through the "grid marking positioning + dynamic parameter adaptation" technology, the problem of inaccurate description of the position of the commodity by the visual large model and the incoordination of the close-up and the picture is solved. The video frame is divided into a 50x50 pixel grid and marked with a red dot with a serial number, the visual large model only needs to output the serial number to accurately position the commodity, avoiding the "illusion phenomenon"; at the same time, the close-up parameters are dynamically adjusted according to the type of the commodity, and the fade-in and fade-out transition is realized to achieve smooth transition of the close-up and the main picture. This process ensures that the close-up of the commodity highlights the details without damaging the integrity of the picture, so that the audience can clearly capture the key features of the commodity, the visual expression is equivalent to manual fine editing, but the efficiency is greatly improved.
[0038] The fifth step is intelligent keyword extraction and material matching: extract keywords and sort them by frequency and importance, generate text descriptions for each material including details, purpose, and emotion, convert keywords and material descriptions into vectors, and by matching materials with high similarity rankings, prioritize matching corresponding dynamic materials when preset keywords appear frequently and insert them into corresponding time points. Specifically, a large language model and precision prompt words are combined to extract keywords: scene text and highlight sentences are input, and the prompt words are set to "extract product name, core buying point, discount information, and emotional tendency". The extracted keywords are sorted according to the frequency of occurrence and preset importance, such as "limited-time discount" has a higher weight than "color". The material is preprocessed to generate a description text according to "detailed description + purpose + emotional label". For example, the sticker "red 'limited-time' label" is described as "a red rounded rectangular label used to highlight limited-time discounts, with an emotional tendency of 'urgent'". The keywords and material text descriptions are converted into vectors, and the matching degree is calculated by cosine similarity. The top three matching materials are selected. If the keyword "discount" appears more than 5 times, dynamic stickers with countdown animations are matched first, and the matching materials are inserted into the corresponding positions according to the timeline. For example, when the sentence corresponding to the keyword "discount" appears, a sticker is added in the upper right corner of the screen.
[0039] The "Language Large Model Keyword Extraction + Material Text Description Vector Matching" mechanism overcomes the limitations of traditional TF-IDF and other algorithms in keyword extraction and the frequent interference in material matching. The language large model, combined with fine-tuned prompt words, extracts precise keywords such as product names and core selling points. The material is then converted into a vector through textual processing using "detailed description + purpose + emotional labeling." Cosine similarity matching is then used to reduce interference from direct image vector matching, increasing the material matching hit rate by over 40%. Furthermore, when pre-set keywords appear frequently, pre-set stickers are prioritized for matching, ensuring a high degree of alignment between the packaging material and the key content of the scene, enhancing the video's message delivery efficiency and visual appeal.
[0040] In the sixth step, the video is integrated, packaged and output adaptively: the main track, the close-up track and the material track content are integrated, the style and length are adjusted according to the output platform, the picture occlusion is detected and adjusted before output, and finally the video and metadata are output. Specifically, the close-up track: the close-up picture of the product is superimposed on the video segment picture corresponding to the main track: the highlight sentence, and then the material track: the matched material is inserted according to the timestamp, and the preset style parameters are adjusted according to the output platform, such as TikTok and Taobao. For example, TikTok adopts fast-paced transition length ≤0.5 seconds and bright color matching; Taobao adopts clear subtitle font size ≥24pt and low saturation sticker. The length is adaptive: if the recommended length of the platform is 15 seconds, the top 5 highlight sentences are retained; if it is 1 minute, the top 10 highlight sentences are retained to ensure the integrity of the core information. Before output, the visual large model is used to detect whether there is picture occlusion, such as sticker occluding the product. If there is occlusion, the material position is automatically adjusted, and finally the video and metadata such as scene title and keyword are output.
[0041] Through the mechanism of "multi-track integration + platform adaptive adjustment + pre-output verification", the problem of single output format and mismatch with platform characteristics in the existing editing method is solved. Multi-track integration ensures the layered cooperation of highlight segments, product close-ups and packaging materials, avoiding picture confusion; the style parameters are adjusted according to the output platform to make the video meet the platform transmission characteristics; the picture occlusion is detected by the visual large model before output and automatically adjusted to ensure the integrity and professionalism of the final video. This process improves the secondary transmission conversion rate of the video on different platforms by more than 30%, and realizes rapid output within 10 minutes after the live broadcast ends, greatly improving the editing efficiency.
[0042] The above-described embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it should not be understood as limiting the scope of the present patent. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present patent should be subject to the appended claims.
Claims
1. A video editing method based on multimodal large model collaboration, characterized in that: The specific steps include: S1. Video preprocessing and intelligent segmentation: Preprocess the original video, cut it into small segments, ensure that each segment ends with a complete sentence, output the timestamp and audio text segment of each segment, and store them in structured data; S2. Multimodal scene segmentation and verification: Each block, i.e., audio text fragment and its timestamp, is input into the language model. The scene is initially segmented according to the rules and then a sprite image is constructed. This is then input into the visual model to verify whether the product has switched. Finally, the scene time, title, and description are output. S3. Product highlight sentence extraction and ranking: The scene text is input into the language model, and the multi-dimensional highlight information is combined to extract the highlight sentences. Audience interaction data and product type preset templates are introduced to adjust the sentence weights, and the sentence is sorted by weight to form a highlight clip timeline. S4. Product visual close-up positioning and insertion: The scene video is divided, labeled, and numbered. The visual model is input to obtain the grid number of the product to locate the product. The close-up parameters are dynamically adjusted according to the product type. The close-up is inserted into the corresponding interval of the highlight sequence through fade-in and fade-out transitions. S5. Intelligent Keyword Extraction and Material Matching: Keywords are extracted and sorted by frequency and importance. Text descriptions are generated for each material, including details, purpose, and emotion. Keywords and material descriptions are converted into vectors. By matching materials with the highest similarity, dynamic materials with high-frequency occurrences of preset keywords are prioritized and inserted at the corresponding time points. S6. Video integration packaging and adaptive output: Integrate the content of the main track, close-up track, and material track, adjust the style and duration according to the output platform, detect and adjust the screen occlusion before output, and finally output the video and metadata.
2. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: The step S1 specifically includes the following steps: S11. Audio feature extraction: Perform audio separation on the original video to extract the sound wave intensity sequence and speech pause features; S12, Semantic Integrity Assessment: Call the speech recognition model to convert the audio into text, analyze the sentence boundaries of the text using the language model, and detect the time points corresponding to punctuation marks, including periods and exclamation marks; S13. Block rule execution: Based on the sound wave intensity, low-intensity segments are prioritized as breakpoints and sentence end times. The video is cut into small segments within a preset time range, ensuring that each segment ends with a complete sentence. S14. Block result storage: Output the start and end timestamps of each block and the corresponding audio text segment, and store them as structured data.
3. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S21. Input each segmented text into the language model, and perform scene segmentation based on preset rules. The following strategies are used for the segmentation results: if no product is segmented or only one product is segmented, it is considered that the product explanation is not completed, and the next subtitle data will be spliced in. If two products are segmented and the second product explanation does not exceed two minutes, it is considered that it has not yet ended. A delay of 2 minutes is added to ensure segmentation stability. If two products are segmented and the second product explanation exceeds two minutes, the first product explanation is considered to be completed, and the first segmentation result is output. The subtitles belonging to the second product are spliced in. If multiple segmentation results appear, only the last two results are retained, and the rest are output. If multiple short segmentation results of less than 2 minutes appear, only the last two results are retained, and the rest are merged into one result and named "complex scene"; S22, extracting frame images from the video at a predetermined time before and after the segmentation point of the preliminary segmentation result at a frequency of one frame per second, and splicing them into sprite images arranged horizontally with a frame spacing of 10 pixels in chronological order; S23. Input the sprite image into the visual model to identify product switching features. If the product switching is visually confirmed, retain the initial classification result. S24: If the segmentation has not been switched, extend the segmentation point backward according to the continuity of the slow image, and re-execute steps S21-S23 until the segmentation is visually confirmed to be accurate; S25. Output the scene text of each scene including the start and end time, scene title, and content introduction.
4. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: The step S3 specifically includes the following steps: S31. Input the scene text into the language model, extract the highlight sentences based on the preset highlight dimension, and output the start and end time and content of the sentences; S32. Introducing auxiliary data to optimize highlight sorting, including combining live comment keywords and peak like times, adding audience interaction weight to sentences mentioning corresponding content, and product type weighting in preset sorting templates based on product type; S33. Sort the highlight sentences in template order according to the weight calculation result to form a basic timeline of the highlight clips.
5. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: When inserting a visual close-up of a product, step S4 divides the video frame into a 50×50 pixel grid, marks a red dot with a black outline in the center of each grid and labels it with a serial number to form a marked frame image, inputs the marked frame image into the visual macro model, outputs the grid serial number of the product in combination with the highlight sentence content, dynamically adjusts the close-up parameters according to the product type, inserts the generated product close-up picture into the time interval corresponding to the highlight sentence, and achieves a smooth transition of the picture through fade-in and fade-out transitions.
6. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: In step S5, when performing intelligent keyword extraction and material matching, a large language model and a precision prompt word combination are used to extract keywords: scene text and highlight sentences are input, and the prompt word is set to "extract product name, core buying point, discount information, emotional tendency". The extracted keywords are sorted according to the frequency of occurrence and preset importance, and the material is preprocessed to generate a description text according to "detailed description + purpose + emotional label". The keywords and material text descriptions are converted into vectors, and the matching degree is calculated by cosine similarity. The materials with the top three matching degrees are selected. If the frequency of occurrence of the preset keywords is greater than the preset frequency, the dynamic stickers of the countdown animation are matched first, and the matching materials are inserted into the corresponding positions according to the timeline.
7. The video editing method based on multimodal large model collaboration according to claim 1 is characterized in that: In step S6, when performing video integration packaging and adaptive output, the close-up track: product close-up image and the main track: video segment images corresponding to the highlight sentence are superimposed, and then the material track: matching materials are inserted according to the timestamp. According to the preset style parameters and duration parameters of the output platform, before output, the visual large model is used to detect whether there is any occlusion in the image. If there is occlusion, the material position is automatically adjusted to output the final video and metadata.
8. A video editing system based on multimodal large model collaboration, characterized by: The system is used to implement the video editing method based on multimodal large model collaboration as described in any one of claims 1 to 7, comprising: a video preprocessing and intelligent segmentation module, a multimodal scene separation and verification module, a product highlight sentence extraction and sorting module, a product visual close-up positioning and insertion module, an intelligent keyword extraction and material matching module, and a video integration packaging and adaptive output module; Video preprocessing and intelligent segmentation module: This module preprocesses the original video, cuts it into small chunks, ensures that each chunk ends with a complete sentence, outputs the timestamp and audio text fragment of each chunk, and stores them in structured data. Multimodal scene segmentation and verification module: This module inputs each segment into the language model, initially divides the scene according to the rules, and then constructs the sprite map. This module inputs the visual model to verify whether the product has been switched, and finally outputs the scene time, title, and description. Product highlight sentence extraction and sorting module: This module is used to input scene text into the language model, extract highlight sentences based on multi-dimensional highlight information, introduce audience interaction data and product type preset templates to adjust sentence weights, and sort by weight to form a highlight clip timeline; Product visual close-up positioning and insertion module: This module is used to divide, mark, and number scene videos, input the visual model to obtain the grid number where the product is located to locate the product, dynamically adjust the close-up parameters based on the product type, and insert the close-up into the corresponding interval of the highlight sequence through fade-in and fade-out transitions; Intelligent keyword extraction and material matching module: This module extracts keywords and sorts them by frequency and importance. It generates a text description of each material, including details, purpose, and emotion. It converts keywords and material descriptions into vectors. By matching materials with the highest similarity, it prioritizes dynamic materials with high-frequency occurrences of preset keywords and inserts them into the corresponding time points. Video integration packaging and adaptive output module: used to integrate the content of the main track, close-up track, and material track, adjust the style and duration according to the output platform, detect and adjust the screen occlusion before output, and finally output the video and metadata.
Citation Information
Patent Citations
Intelligent video editing method and system based on large model
CN117812386A
Video editing method and device and electronic equipment
CN118175400A
Video segmentation method and system based on cross attention and sequence attention
CN118658104A
Video generation method and device, electronic equipment, computer readable storage medium and computer program product
CN119316683A
Video clip label generation method and system for scene change detection
CN119763013A
Cited By
Voice segmentation intelligent editing system based on deep learning
CN121260170A