Highlight combination set generation method and system

The method integrates multi-modal analysis to align and merge video segments coherently, addressing logical disintegration and core information misidentification in smart editing, enhancing viewer engagement.

CN120321474AActive Publication Date: 2025-07-15GUANGZHOU TAIDONG TECH CO LTD

Patent Information

Application Number
CN202510779071.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing intelligent editing technology has problems such as fragment logic fragmentation and inaccurate grasp of core information in the integration of short videos. Especially in multiplayer scenes, it is difficult to identify the interaction weight between the protagonist and the supporting role, resulting in a chaotic rhythm of mixed clipping and a decline in user viewing experience.

Method used

A multimodal model is used to extract highlight candidate fragments, and a high-light collection is generated through space-time constraints and forced alignment of visual content with text summary semantics, including video material segmentation, multimodal feature fusion, timestamp checksum fragment merging, and forced alignment and coherence loss function adjustments are used to use the CLIP model.

Benefits of technology

It realizes accurate extraction of highlight candidate segments and logical coherence in time, improves the core information grasp and audience viewing experience of video integration, and reduces the probability of information redundancy and logical separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321474A_ABST
    Figure CN120321474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a highlight set generation method and system, and the method comprises the steps: obtaining a video material, carrying out the segment segmentation of the video material, inputting the segmented video material into a multi-modal model, carrying out the highlight extraction, obtaining a plurality of highlight candidate segments, and carrying out the highlight extraction based on a preset space-time constraint. The method comprises the following steps of: performing timestamp verification on a plurality of highlight candidate fragments, then performing forced alignment of visual contents and text abstract semantics on the highlight candidate fragments passing the timestamp verification, and finally performing fragment combination on the highlight candidate fragments passing the forced alignment according to a time sequence, and outputting a highlight combination set. Compared with the prior art, the method provided by the invention overcomes the technical problem of inaccurate mastering of integrated video logic splitting and core information in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology. More specifically, the present invention relates to a method and system for generating high-light aggregation. Background Art

[0002] In the digital marketing era, advertising editing technology has become the core means to enhance the communication effect. With the explosive growth of short video platforms and social media, advertising content needs to capture the audience's attention in a very short time. Editing technology significantly improves the conversion rate and memory points of advertisements through rhythm control, narrative logic, and audiovisual impact optimization. The current technology has developed from traditional linear editing to an intelligent editing system that integrates AI, multi-modal analysis, and automated tools.

[0003] The existing intelligent editing technology is mainly implemented by the following several schemes: 1. Editing scheme based on key frame detection: By analyzing the saliency of video frames (such as faces, movements, color contrasts), key frames are extracted as candidate segments. Models such as YOLO and Faster R-CNN are used to identify high-concern areas (such as close-ups of celebrities, goal-scoring moments).

[0004] 2. Editing scheme based on action / motion analysis: Through optical flow analysis, pose estimation, and motion trajectory detection, high-dynamic segments (such as goals in sports events, dance climaxes) are identified. The intensity of the motion is detected, and high-energy segments (such as basketball dunks, racing overtakes) are screened.

[0005] 3. Editing scheme based on emotion / audio analysis: By combining audio waveforms (cheers, climax of background music) and audience reactions (bullet screens, expression recognition), segments with high emotions are extracted.

[0006] 4. Intelligent editing scheme based on multi-modal fusion: By combining multi-modal information such as vision, audio, and text (captions / bullet screens), high-light segments are comprehensively judged.

[0007] However, due to the limitations of the existing model kernels or constraints, the above schemes have the following technical defects: Firstly, the fragment logic is fragmented, making it difficult to support the narrative rhythm of short videos: In sports event or variety show mixed cuts, the existing technology only extracts independent high-light segments (such as shots, funny points), but ignores the causal relationship of events (such as tactical cooperation before a goal). The rhythm of the mixed cut video is chaotic, and it is difficult for users to understand the narrative main line when watching, resulting in a decrease in the completion rate.

[0008] Second, the recognition of character relationships in multi-person scenarios fails, and the mixed editing deviates from the core content: For short videos with multi-person interactions such as variety shows and Vlogs, existing models cannot distinguish the interaction weights between the protagonist and the supporting roles (such as misselecting the close-up of an extra). For example, in a travel mixed edit, the shots of secondary characters are frequently switched, weakening the experience story line of the core characters and reducing user resonance.

[0009] In summary, the existing intelligent editing technology has problems of logical disconnection in integrating videos before and after and inaccurate grasp of core information. Summary of the Invention

[0010] To solve the technical problems of logical disconnection in integrating videos and inaccurate grasp of core information in the existing technology, the present invention discloses a method and system for generating a high-light collection.

[0011] In a first aspect, the method of the present invention discloses a method for generating a high-light collection, including: Obtain video materials; After segmenting the video materials into segments, input them into a multi-modal model for high-light extraction to obtain multiple high-light candidate segments; Based on preset spatio-temporal constraints, perform timestamp verification on multiple high-light candidate segments; Perform forced alignment of the visual content and text summary semantics of the high-light candidate segments that pass the timestamp verification; Merge the high-light candidate segments that pass the forced alignment in chronological order and output a high-light collection.

[0012] Beneficial effects: The method of the present invention extracts high-light candidate segments through a multi-modal model. This method can combine multi-modal information such as vision, audio, and text for fusion analysis, so as to accurately grasp the core information. Through the spatio-temporal constraints pre-configured for the multi-modal model, the timestamp verification of the high-light candidate segments can be automatically performed, and the whole process does not require manual operation, improving the verification efficiency. On this basis, after the high-light candidate segments pass the verification, the method of the present invention also performs forced alignment of the visual content and text summary semantics of the high-light candidate segments to overcome the problem of logical disconnection in integrating videos before and after.

[0013] Preferably, the spatio-temporal constraints include: Check whether the duration of the current high-light candidate segment is greater than or equal to a preset minimum time interval; Verify whether the overlap degree between the current high-light candidate segment and the next high-light candidate segment is less than or equal to an overlap degree threshold; Check whether the significance of the current high-light candidate segment is greater than or equal to a significance threshold; If all the above conditions are satisfied, the current high-light candidate segment is determined to pass the timestamp verification; if any of the above conditions is not satisfied, the current high-light candidate segment is determined not to pass the timestamp verification.

[0014] Beneficial effects: When verifying the timestamp of the high-light candidate segment, by verifying the overlap degree between the current high-light candidate segment and the next high-light candidate segment, the overlap degree between adjacent high-light candidate segments in time can be reasonably regulated, reducing the probability of too high or too low overlap degree. Among them, too high overlap degree will lead to information redundancy, while too low overlap degree will lead to disconnection of information before and after, exacerbating the logical disconnection between the front and back of the video. And through the above method, the problem of information redundancy can be overcome and the problem of logical disconnection can be solved. In addition, the above spatio-temporal constraint introduces the judgment of saliency, which can screen out the core information more quickly, thus further overcoming the problem of inaccurate grasp of the core information in the prior art.

[0015] Further, the expression of the spatio-temporal constraint is:

[0016] In the formula, represents the timestamp verification condition depending on the start time and the end time ; represents the minimum time interval that meets the timestamp verification condition; represents the overlap degree operation function; represents calculating the overlap degree between the current high-light candidate segment and the next high-light candidate segment ; represents the overlap degree threshold; represents calculating the saliency of the current high-light candidate segment ; represents the preset training coefficient; represents calculating the maximum saliency value of all high-light candidate segments in the video material .

[0017] Preferably, if all high-light candidate segments do not pass the timestamp verification or the number of high-light candidate segments passing the timestamp verification does not reach the expected number, a preset backtracking mechanism is triggered, and the multi-modal model is called to re-extract the high-light candidate segments.

[0018] Preferably, after outputting the high-light set, the method of the present invention further includes: scoring the merging quality of the high-light set to obtain an evaluation value; if the evaluation value is greater than or equal to the expected score, taking the high-light set as the final integrated video; if the evaluation value is less than the expected score, returning to the high-light extraction step.

[0019] Beneficial effects: After obtaining a high photosynthetic set, the method of the present invention will also evaluate the merging quality of the high photosynthetic set. If the merging quality of the high photosynthetic set does not meet the expectation, it means that the information extracted by high light is not important enough. Then, it returns to the high light extraction step and performs iterative extraction again, thereby further solving the problem of inaccurate grasp of core information in the prior art.

[0020] Preferably, the merging quality of the high photosynthetic set is scored, and the expression for obtaining the evaluation value is:

[0021] In the formula, represents the evaluation value, represents the MLLM (Multimodal Large Language) model, represents the high photosynthetic set, represents the expected information of quality evaluation.

[0022] Preferably, the expected information of quality evaluation includes rhythm expected parameters, plot logic expected descriptions, and attraction expected descriptions.

[0023] Preferably, the forced alignment of the visual content and the text summary semantics of the high light candidate fragments passed through the timestamp verification includes: Using a preset CLIP (Contrastive Language–Image Pre-training) model, the forced alignment of the visual content and the text summary semantics of the high light candidate fragments passed through the timestamp verification is performed.

[0024] Preferably, the CLIP model is set with a coherence loss function for guiding model training. The coherence loss function is specifically:

[0025] In the formula, represents the alignment loss value, represents from to the summation symbol, represents the total number of segmented fragments, represents the calculation function of the CLIP model for forced alignment of the visual content in the th fragment, represents the calculation function of the CLIP model for forced alignment of the text summary semantics in the th fragment for forced alignment, represents the L2 norm.

[0026] Beneficial effects: The method of the present invention designs a set of coherence loss functions to guide the model to adjust parameters, so that the difference between video features and text features gradually decreases.

[0027] In a second aspect, the present invention also discloses a high - photosynthesis collection generation system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the high - photosynthesis collection generation method described in the first aspect is implemented.

[0028] The beneficial effects of the present invention are as follows: 1. Compared with the prior art, the method of the present invention extracts high - light candidate segments through a multi - modal model. This method can combine multi - modal information such as vision, audio, and text for fusion analysis, so as to accurately grasp the core information. After the high - light candidate segments pass the verification, the method of the present invention also performs forced alignment of the visual content and the text summary semantics of the high - light candidate segments to overcome the problem of logical disconnection before and after integrating the video.

[0029] 2. Compared with the prior art, the multi - modal model of the method of the present invention can automatically generate spatio - temporal constraints according to the time - stamp verification rules and reasonably regulate the overlap degree between adjacent high - light candidate segments in time, which not only solves the problem of information redundancy but also further solves the problem of logical disconnection.

[0030] 3. Compared with the prior art, the method of the present invention designs a set of coherence loss functions to guide the model to adjust parameters, so that the difference between video features and text features gradually decreases. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] By reading the following detailed description with reference to the accompanying drawings, the above - mentioned and other objects, features, and advantages of the exemplary embodiments of the present invention will become easy to understand. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where: Figure 1 is a flowchart of the high - photosynthesis collection generation method in the first embodiment of the method of the present invention; Figure 2 is a schematic structural diagram of the high - photosynthesis collection generation system in the third embodiment of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0033] Next, the specific implementation manners of the present invention will be described in detail with reference to the accompanying drawings.

[0034] Embodiment 1 As Figure 1 shown, this embodiment discloses a high-light synthesis method, including: S10: Obtain video materials.

[0035] In this embodiment, the video materials can be any one of competition videos, TV dramas, advertising and promotion videos, public welfare promotion videos or online recorded videos. The video materials can be imported into the computer system corresponding to the method of the present invention by professionals, or automatically obtained by the computer system according to the preset action response time.

[0036] S20: Segment the video materials and then input them into a multi-modal model for high-light extraction to obtain multiple high-light candidate segments.

[0037] In this embodiment, during the process of segmenting, there will be partial content overlap between two adjacent segmented segments in the time series. The above multi-modal model can adopt the VideoLLM multi-modal model, which is a high-precision extraction model capable of performing visual feature extraction, sound source feature extraction and text feature extraction, and fusing the above features.

[0038] Specifically, extract the multi-modal fusion features of the segmented segments through the VideoLLM multi-modal model, analyze and calculate the multi-modal fusion features, and then use the segmented segments corresponding to the analyzed and calculated multi-modal fusion features as high-light candidate segments.

[0039] S30: Based on the preset spatio-temporal constraints, perform timestamp verification on multiple high-light candidate segments.

[0040] S40: Force-align the visual content and text summary semantics of the high-light candidate segments that pass the timestamp verification.

[0041] In this embodiment, the model participating in the forced alignment can adopt the CLIP model.

[0042] It should be explained that the CLIP (Contrastive Language–Image Pre-training) model is a contrastive language-image pre-training model developed by OpenAI. It aims to learn the semantic associations between images and texts by pre-training on a large number of image-text pairs, enabling the model to understand the semantic content of images and texts and achieve mutual matching and retrieval between images and texts.

[0043] S50: Merge the high-light candidate segments that pass the forced alignment in chronological order and output the high-light synthesis.

[0044] Through the above technical solution, the method of the present invention can perform fusion analysis by combining multi-modal information such as vision, audio, and text through step S20, and accurately grasp the core information. After the highlight candidate segment passes the verification, the method of the present invention can also perform forced alignment of the visual content and text summary semantics of the highlight candidate segment through step S40 to overcome the problem of logical disconnection before and after integrating the video.

[0045] Compared with the prior art, the method of the present invention grasps the core information more accurately, and the generated highlight compilation is logically coherent in time, improving the viewing experience of the audience.

[0046] Embodiment 2 On the basis of Embodiment 1, this embodiment mainly gives a further explanatory description of the method of the present invention at the core algorithm level.

[0047] Exemplarily, in step S20, the algorithm expression for obtaining the fusion feature of the above VideoLLM multi-modal model is:

[0048] In the formula, represents the th multi-modal fusion feature of the video material, represents the multi-modal large model, represents the th segmented segment of the video material, represents the cross-modal attention fusion operator, represents the visual modality feature of the th segmented segment in the video material, represents the audio modality feature of the th segmented segment in the video material, represents the text or subtitle modality feature of the th segmented segment in the video material.

[0049] Exemplarily, in the process of real-time analysis and processing of the multi-modal fusion feature, it is necessary to extract the historical segment feature and the current segment feature from the multi-modal fusion feature, and score the currently processed segmented segment. The specific algorithm expression is:

[0050] In the formula, represents the score weight matrix; represents the historical segment feature; represents the current segment feature; represents the query vector, which is generated based on the information of the current time step and is used to query in the historical and current segment features; It represents the concatenation operation of historical segment features and current segment features to form a more comprehensive feature representation; It represents the dimension of the feature. When the weight score corresponding to the segmented segment reaches the expected value, it indicates that the segmented segment belongs to the core information, and this segmented segment is used as a highlight candidate segment.

[0051] Furthermore, the spatio-temporal constraint in step S20 of the method of the present invention is specifically as follows: Condition 1: Check whether the duration of the current highlight candidate segment is greater than or equal to a preset minimum time interval; Condition 2: Check whether the overlap degree between the current highlight candidate segment and the next highlight candidate segment is less than or equal to the overlap degree threshold.

[0052] Condition 3: Check whether the significance of the current highlight candidate segment is greater than or equal to the significance threshold.

[0053] If all of the above Conditions 1 to 3 are satisfied, the current highlight candidate segment is determined to pass the timestamp verification. If any of the above conditions is not satisfied, the current highlight candidate segment is determined to fail the timestamp verification.

[0054] Specifically, the specific expression of the above spatio-temporal constraint is:

[0055] In the formula, represents the timestamp verification condition depending on the start time and the end time of the timestamp verification condition, represents the minimum time interval that meets the timestamp verification condition, represents the overlap degree operation function, represents calculating the current highlight candidate segment and the next highlight candidate segment for the overlap degree, represents the overlap degree threshold, represents calculating the significance of the current highlight candidate segment of the current highlight candidate segment, represents the preset training coefficient, represents calculating the maximum value of the significance of all highlight candidate segments in the video material of the video material.

[0056] Through the above design of spatio-temporal constraints, Condition 2 can reasonably regulate the overlap degree between temporally adjacent highlight candidate segments by verifying the overlap degree between the current highlight candidate segment and the next highlight candidate segment, reducing the probability of the overlap degree being too high or too low. More specifically, too high an overlap degree will lead to information redundancy and resource waste, thereby reducing the data processing efficiency. While too low an overlap degree will lead to discontinuous information before and after, thereby exacerbating the logical fragmentation of the video before and after. And verifying through the above Condition 2 can overcome both the problem of information redundancy and the problem of logical fragmentation. In addition, Condition 3 in the above spatio-temporal constraints also introduces a saliency judgment, which can quickly screen out the core information, thereby further overcoming the problem of inaccurate grasp of the core information in the prior art.

[0057] Furthermore, if all highlight candidate segments fail the timestamp verification or the number of highlight candidate segments passing the timestamp verification does not reach the expected number, it indicates that the coherence of the context logic of the highlight candidate segments extracted by the multi-modal model does not meet the requirements. At this time, it is necessary to trigger the backtracking mechanism, return to step S20, re-call the multi-modal model, and re-extract the highlight candidate segments.

[0058] Furthermore, the model for forcibly aligning the visual content and text summary semantics of the highlight candidate segments in the above step S40 adopts the CLIP model. For this model, a more targeted coherence loss function is designed in this embodiment, specifically:

[0059] In the formula, represents the alignment loss value, represents the summation symbol from to ; represents the total number of segmented segments, represents the calculation function of the CLIP model for forcibly aligning the visual content in the th segment, represents the calculation function of the CLIP model for forcibly aligning the text summary semantics in the th segment, represents the L2 norm.

[0060] It should be noted that the above coherence loss function can quantify the semantic differences between video features and text features. The larger the alignment loss value, the greater the semantic differences between the two and the lower the alignment degree; the smaller the alignment loss value, the smaller the semantic differences between the two and the higher the alignment degree. The CLIP model uses the alignment loss value as the judgment basis to verify the semantic differences between video features and text features. If the alignment loss value is greater than or equal to the alignment expected value, it indicates that the forced alignment is successful, and the high-light candidate segments with successful alignment enter the data processing flow of step S50. If the alignment loss value is less than the alignment expected value, it indicates that the forced alignment fails. The hyperparameters in the CLIP model are re-adjusted and the forced alignment is performed again. If the forced alignment fails multiple times, the high-light candidate segment is discarded.

[0061] Compared with the prior art, the above coherence loss function can play a role in guiding the CLIP model to adjust parameters, so that the differences between video features and text features gradually decrease during multiple iterative trainings.

[0062] Furthermore, after step S50, the method of the present invention further includes: S60: Score the merging quality of the high-light set to obtain an evaluation value.

[0063] S70: If the evaluation value is greater than or equal to the expected score, use the high-light set as the final integrated video.

[0064] S80: If the evaluation value is less than the expected score, return to step S20.

[0065] Specifically, the calculation expression of the above step S60 is:

[0066] In the formula, represents the evaluation value, represents the MLLM (Multimodal Large Language) model, represents the high-light set, represents the quality evaluation expected information.

[0067] In this embodiment, the quality evaluation expected information includes rhythm expected parameters, plot logic expected descriptions, and attraction expected descriptions.

[0068] It should be noted that in this embodiment, the quality assessment expected information can be input by the user or the information in the pre-set database can be called. The rhythm expected parameters include specific parameters such as video beat parameters, video multiples, and climax time domains of the plot. The description of the plot logic expectation is a relatively vague language description. These descriptions of the plot logic expectation can be "starting with depression and then rising", "starting with parts and then summarizing", "starting with summarization and then parts", "describing with a single character (single-image plot)", or "describing with multiple characters separately and then presenting the characters in the same frame with a general plot segment (group-image plot)", "flashback plot", "intercalary plot", or "sequential plot", etc., related or similar language descriptions. The description of the attraction expectation is also a relatively vague language description. These descriptions of the attraction expectation can be "highlighting popular traffic stars, internet celebrities, or influencers", or "highlighting current hot topics, popular memes, or popular traffic topics", or "highlighting current popular electronic, automotive, or sports products". These vague language descriptions will be parsed by the language understanding model, and then keywords or sentences will be extracted. Then, an online search of network information will be performed on the online platform to calculate the traffic degree of the keywords or sentences. Then, the traffic degree in this quantitative way will be fused with the rhythm expected parameters as the quality assessment expected information, so as to achieve multi-dimensional scoring of the high-light collection.

[0069] Through the above steps S60 - S80, after obtaining the high-light collection, the method of the present invention will also evaluate the merging quality of the high-light collection. If the merging quality of the high-light collection does not meet the expectation, it means that the information extracted by the high-light extraction is not important enough. Return to step S20 to iteratively extract the high-light candidate segments that meet the expectation again, thereby further improving the accuracy of the method of the present invention in grasping the core information.

[0070] Embodiment III As Figure 2 shown, this embodiment discloses a high-light collection generation system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the high-light collection generation method described in Embodiment I or Embodiment II is implemented.

[0071] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

Claims

1. A high photosynthesis integration generation method, characterized in that, Including: Obtain video materials; After segmenting the video materials, input them into a multi-modal model for highlight extraction to obtain multiple highlight candidate segments; Based on preset spatio-temporal constraints, perform timestamp verification on multiple highlight candidate segments; Perform forced alignment of the visual content and text summary semantics of the highlight candidate segments that pass the timestamp verification; Merge the highlight candidate segments that pass the forced alignment in chronological order and output a highlight collection.

2. The high photosynthetic set generation method according to claim 1, characterized in that The spatio-temporal constraints include: Check whether the duration of the current highlight candidate segment is greater than or equal to a preset minimum time interval; Verify whether the overlap between the current highlight candidate segment and the next highlight candidate segment is less than or equal to the overlap threshold; Check whether the significance of the current highlight candidate segment is greater than or equal to the significance threshold; If all the above conditions are met, the current highlight candidate segment is determined to pass the timestamp verification; if any of the above conditions is not met, the current highlight candidate segment is determined to fail the timestamp verification.

3. The high photosynthetic set generation method according to claim 2, characterized in that The expression of the spatio-temporal constraints is: In the formula, represents a timestamp verification condition that depends on the start time and the end time ; represents the minimum time interval that meets the timestamp verification condition, represents an overlap degree operation function, represents calculating the overlap degree between the current highlight candidate segment and the next highlight candidate segment, represents an overlap degree threshold, represents calculating the significance of the current highlight candidate segment, represents a preset training coefficient, represents calculating the maximum value of the significance of all highlight candidate segments in the video material.

4. The high photosynthetic collection generation method according to claim 1, characterized in that If all highlight candidate segments fail the timestamp verification or the number of highlight candidate segments that pass the timestamp verification does not reach the expected number, trigger a preset backtracking mechanism, call the multi-modal model, and re-extract the highlight candidate segments.

5. The high photosynthetic collection generation method according to claim 1, wherein After outputting the highlight collection, the method further includes: Score the merging quality of the highlight collection to obtain an evaluation value; If the evaluation value is greater than or equal to the expected score, use the highlight collection as the final integrated video; If the evaluation value is less than the expected score, return to the highlight extraction step.

6. The high photosynthetic collection generation method according to claim 5, characterized in that The expression for scoring the merging quality of the highlight collection to obtain an evaluation value is: In the formula, represents the evaluation value, represents the MLLM model, represents the high photosynthetic set, represents the quality assessment expected information.

7. The high photosynthetic collection generation method according to claim 6, characterized in that The quality evaluation expected information includes rhythm expected parameters, plot logic expected descriptions, and attraction expected descriptions.

8. The high photosynthetic collection generation method according to claim 1, characterized in that Performing forced alignment of the visual content and text summary semantics of the highlight candidate segments that pass the timestamp verification includes: Using a preset CLIP model to perform forced alignment of the visual content and text summary semantics of the highlight candidate segments that pass the timestamp verification.

9. The high photosynthetic set generation method according to claim 8, characterized in that, The CLIP model is set with a coherence loss function for guiding model training, and the coherence loss function is specifically: Wherein, represents the alignment loss value, represents the sum symbol from to ; represents the total number of segmentation segments, represents the calculation function for the CLIP model to force alignment of the visual content in the -th segment, represents the calculation function for the CLIP model to force alignment of the text summary semantics in the -th segment, ; represents the L2 norm.

10. A high photosynthetic collection generation system, characterized in that, Including a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method for generating a highlight collection according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Automatic generation method of video selection set based on content analysis

    CN112445935A

  • A highlight fragment extraction method and device, computer equipment and a storage medium

    CN113408461A

  • Live video key point marking method and device, equipment and storage medium

    CN119071520A

  • Video highlight detection method based on weak supervision multi-mode large model

    CN119785257A

  • Vision-text collaborative abstract generation method and system based on multi-modal learning

    CN119862861A

Cited By

  • Video editing method and system based on multi-modal large model collaboration

    CN120786154A