A high light point extraction method and system based on a multi-modal model
By using a multimodal model and a spatiotemporal cross-attention mechanism to filter highlight points, the problem of low efficiency in extracting advertising video materials is solved, achieving efficient and accurate highlight point extraction and improved user experience.
Patent Information
- Application Number
- CN202510807807.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing technologies for highlight extraction in the advertising field suffer from poor targeting and low efficiency. Traditional methods struggle to identify highlights in advertising video materials, and the inference engine of the black-box model requires multiple training sessions during transfer, leading to reduced efficiency.
A highlight extraction method based on a multimodal model is adopted. By segmenting video footage into overlapping sub-segments, joint feature extraction is performed. Visual, audio, and text modal features are combined, and an attention score matrix is calculated using a spatiotemporal cross-attention mechanism to filter highlight points. A highlight inference engine and a multimodal reward function are configured to improve model performance.
It achieves a deep understanding of advertising video materials, enhances the ability to express features and extract targeted features, improves the accuracy of focusing and positioning of highlights, and enhances the user viewing experience.
Smart Images

Figure CN120708124B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image data extraction. More particularly, the present application relates to a highlight point extraction method and system based on a multi-modal model. BACKGROUND
[0002] With the explosive growth of short video platforms and social media, advertising content often needs to capture the attention of the audience in a very short time, and the few pictures that can capture the audience's attention in a very short time are often visually powerful or sensory empathetic. In the process of traditional advertisement picture extraction, professional personnel often extract high light points from video materials as high light images or segments for advertising and promotion.
[0003] The above-mentioned traditional highlight point extraction method is time-consuming and labor-intensive. In order to improve efficiency, CN119763014A discloses a video content analysis method, device, medium and product based on deep learning, which mainly relies on a spatio-temporal pyramid attention network and a dynamic contrast learning framework to extract key frame images. This method has the following disadvantages if directly applied to the field of advertising to extract high light segments:
[0004] 1. Lack of targeting, difficult to identify high light points of advertising video materials.
[0005] 2. The inference engine uses a black box model, which cannot optimize the performance of the spatio-temporal pyramid attention network. When migrating to advertising video analysis, it is difficult to set the inference rule engine, resulting in an increase in the number of training rounds and a decrease in the overall efficiency of high light point extraction.
[0006] Therefore, the prior art has the problems of poor targeting and low extraction efficiency in the application of the advertising field. SUMMARY
[0007] To solve the above-mentioned technical problems of poor targeting and low extraction efficiency of the prior art, the present application discloses a highlight point extraction method and system based on a multi-modal model.
[0008] In a first aspect, the present application discloses a highlight point extraction method based on a multi-modal model, comprising:
[0009] In response to the input of the video material, the video material is segmented into segments to obtain a plurality of overlapping sub-segments;
[0010] The plurality of overlapping sub-segments are input into a pre-set multi-modal large model for joint feature extraction to obtain mixed features;
[0011] The historical segment features and the current segment features are extracted from the mixed features;
[0012] The historical segment feature and the current segment feature are input into a preset spatio-temporal cross attention mechanism model, and a score matrix of attention in a time dimension is calculated.
[0013] The highlight segment or highlight node corresponding to the element with a score greater than a preset threshold in the score matrix of attention is taken as a highlight highlight.
[0014] Beneficial effects: For the extraction of highlight highlights of video materials in the advertising field, the method of the application can deeply understand the video materials by using a multi-modal large model, realize the complementation of different modal information, enhance the expression ability of features and the pertinence of feature extraction, and on this basis, the method of the application uses a spatio-temporal cross attention mechanism model to calculate the attention score matrix of the features, and screens the highlight highlights more efficiently based on the attention score matrix, so as to realize accurate focusing and positioning in the time dimension and association with the context. Compared with the prior art, the method of the application solves the technical problems of poor pertinence and low extraction efficiency of the prior art.
[0015] Preferably, the plurality of overlapping sub-segments are input into a preset multi-modal large model for joint feature extraction, and the specific expression of the mixed features is:
[0016]
[0017] In the formula, H i represents the i-th mixed feature of the video material, VideoLLM represents the multi-modal large model, S i represents the i-th overlapping sub-segment of the video material, represents a cross-modal attention fusion operator, f vision (V i ) represents the i-th visual modal feature in the video material, f audio (A i ) represents the i-th audio modal feature in the video material, f text (T i ) represents the i-th text or subtitle modal feature in the video material.
[0018] Beneficial effects: For the extraction of mixed features, the method of the application extracts features from three modalities of visual modalities, audio modalities and subtitle modalities, compared with traditional single-modal feature extraction (mainly referring to visual modalities), the method of the application fully considers the particularity of advertising video materials (i.e. on the basis of video performance, exaggerated or inductive voice emphasis or artistic text will be used), and can more accurately and pertinently extract effective features.
[0019] Preferably, before the plurality of overlapping sub-segments are input into the preset multi-modal large model for joint feature extraction to obtain the mixed features, the method of the application further comprises:
[0020] Define the fragment overlap rate for multiple overlapping sub-fragments;
[0021] Based on the overlap rate, set long-range dependency constraints for adjacent overlapping sub-segments.
[0022] Preferably, the expression for the long-range dependency constraint is:
[0023] S i+1 ∩S i =α·|S i |
[0024] In the formula, α represents the overlap rate of the segments, and S i+1 S represents the (i+1)th overlapping sub-segment of the video footage. i+1 For S i Temporally adjacent overlapping sub-slices.
[0025] Beneficial effects: Through the design of the long-range dependency constraints described above, the method of this invention can carry a certain amount of contextual information when processing each sub-segment. This allows for better capture of long-range dependencies in the data during subsequent analysis, modeling, and other operations.
[0026] Preferably, before inputting multiple overlapping sub-segments into a preset multimodal large model for joint feature extraction to obtain hybrid features, the method of the present invention further includes:
[0027] Configure a specular inference engine, which is used to drive a multimodal large model to focus on the narrative continuity of video footage.
[0028] Furthermore, the high-intensity inference engine includes a Prompt dynamic evolution mechanism, the specific expression of which is:
[0029]
[0030] In the formula, P k Let P represent the Prompt generated in the k-th iteration. base This represents the basic Prompt; Concat represents the concatenation of functions. This indicates the summary characteristics of the previous reasoning. Used to summarize key information from the multimodal data during the previous inference process.
[0031] Beneficial effects: In order to improve the data processing capability of multimodal large models, the method of the present invention configures a high-light inference engine for multimodal large models. This engine can effectively guide the model to focus on key information and improve the performance of the model.
[0032] Preferably, the highlight inference engine further sets a multi-modal reward function for defining the dual-objective optimization, wherein the defined dual-objective optimization comprises a continuity reward and an audience appeal reward of adjacent overlapping sub-clips.
[0033] Beneficial effects: The method of the present application sets two target optimizations, namely continuity reward and audience appeal reward, which can enhance the content quality of highlight bright spots and improve the viewing experience of users.
[0034] Further, the definition of the multi-modal reward function is:
[0035]
[0036] In the formula, represents a maximum operation on the parameter θ, represents an expectation operator, represents a summation from time step t = 1 to T, and γ t represents a discount factor, and R cohere (S t ) represents a continuity reward of overlapping sub-clips in the state at time step t, λ represents a hyperparameter, and R impact (s t ) represents an audience appeal reward of overlapping sub-clips in the state at time step t.
[0037] Preferably, the spatio-temporal cross-attention mechanism model adopts any one of a decomposed attention framework, a supervised learning framework, or a unified attention framework.
[0038] In a second aspect, the present application further discloses a highlight bright spot extraction system based on a multi-modal model, comprising a processor and a memory, and the memory stores computer program instructions, which realize the highlight bright spot extraction method based on the multi-modal model as recorded in the first aspect when executed by the processor.
[0039] The present application has the following beneficial effects:
[0040] 1. For the extraction of highlight bright spots of video materials in the advertising field, the method of the present application uses a multi-modal large model to deeply understand the video materials, realizes the complementation of different modal information, enhances the expression ability of features and the pertinence of feature extraction; on this basis, the method of the present application uses a spatio-temporal cross-attention mechanism model to calculate an attention score matrix of features, and screens highlight bright spots more efficiently based on the attention score matrix, so as to realize accurate focusing, positioning in the time dimension, and association with the context. Compared with the prior art, the method of the present application solves the technical problems of poor pertinence and low extraction efficiency of the prior art.
[0041] 2. Compared with the prior art, the method of the present application configures a high light reasoning engine for a multi-modal large model, which can effectively guide the model to focus on key information and improve the performance of the model.
[0042] 3. Compared with the prior art, the method of the present application sets up two target optimization directions, namely continuity reward and audience attraction reward, which can enhance the content quality of high light highlights and improve the user viewing experience. BRIEF DESCRIPTION OF DRAWINGS
[0043] The above and other objects, features and advantages of the exemplary embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which several embodiments of the present application are shown by way of example, and wherein like or corresponding elements refer to like or corresponding parts throughout the several views, and in which:
[0044] Figure 1 is a flowchart of the high light highlight extraction method based on a multi-modal model in embodiment one of the present application;
[0045] Figure 2 is a structural schematic diagram of the high light highlight extraction system based on a multi-modal model in embodiment three of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0047] The specific embodiments of the present application will be described in detail below with reference to the drawings.
[0048] Embodiment one
[0049] As shown in the figure, the present embodiment discloses a high light highlight extraction method based on a multi-modal model, which comprises: Figure 1
[0050] S10: In response to the input of the video material, the video material is segmented into multiple overlapping sub-clips.
[0051] In the present embodiment, the video material mainly refers to the video material in the field of advertising, which can also be a popular short drama, a product promotion video, a public service advertisement video, a variety show video or a star video. The overlapping sub-clip refers to a sub-clip with partial content repetition between multiple selected sub-clips in the original video sequence data (i.e. multiple video clips with content intersection between clips).
[0052] S20: inputting the plurality of overlapping subsegments into a preset multi-modal large model to perform joint feature extraction, and obtaining mixed features.
[0053] In this embodiment, the multi-modal large model extracts mixed features including visual modal features, audio modal features, and text or subtitle modal features.
[0054] S30: extracting historical segment features and current segment features from the mixed features.
[0055] S40: inputting the historical segment features and the current segment features into a preset spatio-temporal cross-attention mechanism model to calculate an attention score matrix in the time dimension.
[0056] In this embodiment, the spatio-temporal cross-attention mechanism model is a spatio-temporal cross-attention mechanism model based on a unified attention framework. In other embodiments, the spatio-temporal cross-attention mechanism model can also be a spatio-temporal cross-attention mechanism model based on a decomposed attention framework or a supervised learning framework.
[0057] S50: taking a highlight segment or a highlight node corresponding to an element with a score greater than a preset threshold in the attention score matrix as a highlight highlight.
[0058] It should be explained that a video key frame refers to image data focusing on the structure and representation of video content, while a highlight highlight is a term in the advertising field, which mainly refers to a climax image or segment focusing on the audience's experience. The attention score matrix is a spatio-temporal cross-attention weight matrix, and each element score in the weight matrix represents a weight coefficient of the element as a highlight highlight. The higher the score, the higher the possibility that the video segment or node corresponding to the element can be a highlight highlight. The highlight highlight obtained through the above method can be a momentary static image (highlight node) or a dynamic image with a relatively short time length (highlight segment). The plurality of highlight highlights obtained through the above method are provided as a candidate set to professionals or professional software for editing and integration,
[0059] Through the above steps S10-S50, the method of the present application uses a multi-modal large model to deeply understand video materials, realizes information complementation of visual modal, audio modal, and text modal, and thus enhances the expression ability of features and the pertinence of feature extraction. In addition, the method of the present application uses a spatio-temporal cross-attention mechanism model to calculate an attention score matrix of features, and screens highlight highlights based on the attention score matrix, so as to realize accurate focusing and positioning of highlight highlights in the time dimension and association with the context.
[0060] Through the above technical solutions, the method of the present application solves the technical problems of poor pertinence and low extraction efficiency in the prior art.
[0061] Embodiment Two
[0062] Based on the disclosure of Embodiment One, this embodiment further explains the specific algorithm used in the method of the application.
[0063] In this embodiment, multiple overlapping sub-clips are input into a preset multi-modal large model for joint feature extraction, and the specific expression of the mixed feature is:
[0064]
[0065] In the formula, H i represents the i-th mixed feature of the video material, VideoLLM represents the multi-modal large model, S i represents the i-th overlapping sub-clip of the video material, represents a cross-modal attention fusion operator, f vision (V i ) represents the i-th visual modal feature in the video material, f audio (A i ) represents the i-th audio modal feature in the video material, f text (T i ) represents the i-th text or subtitle modal feature in the video material.
[0066] Preferably, represents a real number space of d dimensions.
[0067] It needs to be explained that in the field of advertising, there are climax parts, inducing parts or exaggerated emphasized parts in the video material. These contents will be embodied in different forms, for example: in order to provide visual stimulation to the audience, the character actions or scene changes in the video material often appear unexpected visual fragments; in order to highlight the performance of the product for sale, some subtitles or bullet screens are often added in the video material; in order to enhance the audience's empathy, the audio part of the video material will be adjusted in volume or dubbing. The above parts may become highlights, and the modal features of these highlights are often multi-modal.
[0068] Through the above algorithm design, the method of the application adopts VideoLLM as a multi-modal large model, which can realize multi-modal feature extraction of the video material, and perform attention fusion operation on the above multi-modal features to obtain more representative mixed features, realizing multi-modal sub-clip encoding.
[0069] Preferably, before inputting multiple overlapping sub-clips into a preset multi-modal large model for joint feature extraction to obtain mixed features, the method of the application further comprises:
[0070] a slice overlap rate defining a plurality of overlapping sub-segments;
[0071] According to the slice overlap rate, a long-range dependency constraint of adjacent overlapping sub-segments is set.
[0072] Specifically, the expression of the long-range dependency constraint is as follows:
[0073] S i+1 ∩S i =α·|S i |
[0074] In the formula, α represents the slice overlap rate, S i+1 represents the i+1th overlapping sub-segment of the video material, S i+1 is the adjacent overlapping sub-segment adjacent in time sequence. i
[0075] In the embodiment, the slice overlap rate is 0.2.
[0076] Through the long-range dependency constraint, the effect of setting the context sliding window can be achieved. In this way, when processing each sub-segment, a certain amount of previous information can be carried. In subsequent analysis, modeling, and other operations, the model or algorithm can better capture the long-range dependency relationship in the data.
[0077] Further, before inputting the plurality of overlapping sub-segments into the preset multi-modal large model for joint feature extraction to obtain mixed features, the method further comprises:
[0078] configuring a highlight reasoning engine for driving the multi-modal large model to focus on the continuity of the plot of the video material.
[0079] Specifically, the highlight reasoning engine includes a Prompt dynamic evolution mechanism and a multi-modal reward function.
[0080] The specific expression of the Prompt dynamic evolution mechanism is as follows:
[0081]
[0082] In the formula, P k represents the Prompt generated in the kth iteration, P base represents the base Prompt, Concat represents the concatenation function, represents the summary feature of the previous reasoning, for summarizing the key information of the multi-modal data in the previous reasoning process.
[0083] It should be explained that, in the embodiment, the Prompt refers to the input text provided to the artificial intelligence model, which is mainly used to guide the model to generate a specific output.
[0084] In the direction of driving multi-modal large models, the above Prompt dynamic evolution mechanism has the following advantages:
[0085] First, it can accurately guide model derivation. Multi-modal large models have the ability to process text, images, and audio, but when faced with complex tasks, they need to be accurately guided to call the deep capabilities of different modalities. The Prompt dynamic evolution mechanism can dynamically adjust the structure and content of the input prompt according to task requirements.
[0086] Second, it can enhance the adaptability of multi-modal large models to different tasks. Different tasks have significant differences in the demand for multi-modal data, and the Prompt dynamic evolution mechanism can adaptively generate more suitable prompts by analyzing task characteristics.
[0087] Third, it can improve the interaction efficiency between users and multi-modal large models. The Prompt dynamic evolution mechanism can understand user intentions in real time, dynamically adjust prompts, and reduce the number of times users need to modify inputs. In the field of intelligent design, after designers propose initial ideas, the mechanism can dynamically optimize prompts based on the text descriptions or sketches input by the designers, guide the model to quickly extract multiple high-light points that meet the requirements, speed up the design iteration process, and improve interaction efficiency.
[0088] Fourth, it can optimize the fusion effect of multi-modal large models. The core of multi-modal large models is the fusion of different modal data, and the Prompt dynamic evolution mechanism can promote the effective integration of information from each modality by optimizing prompts. For example, in the process of overlapping sub-fragment understanding, the mechanism can dynamically generate prompts based on the visual content, audio features, and caption text of the overlapping sub-fragment, guide the model to mine associated information between modalities, such as determining character emotions through voice tone and understanding plot development by combining visual scenes, achieving deeper multi-modal fusion, and improving the model's understanding ability of complex scenes.
[0089] Among them, the specific definition of the multi-modal reward function is:
[0090]
[0091] In the formula, denotes the maximum operation on the parameter θ, denotes the expectation operator, denotes the sum from time step t = 1 to T, γ t denotes the discount factor, R cohere (s t ) denotes the continuity reward of the overlapping sub-fragment in the state at time step t, λ denotes the hyperparameter, R impact (s trepresents the audience appeal reward of the overlapping sub-fragments in the state at time step t.
[0092] In the embodiment, R cohere is calculated by the temporal graph convolution network TGCN, and R impact is predicted based on the audience attention heat map.
[0093] Through the design of the above-mentioned multi-modal reward function, the method has the following advantages:
[0094] First, the performance of the multi-modal large model is further improved. The continuity reward in the reward function can encourage the multi-modal large model to generate content that is logically and semantically consistent in the state at time step t. The audience appeal reward can encourage the multi-modal large model to extract more attractive highlights, thereby meeting the user's demand for extracting high-quality and interesting content.
[0095] Second, the multi-modal fusion capability of the multi-modal large model is further improved. Specifically, the reward function considers rewards in terms of continuity and appeal, guiding the multi-modal large model to better integrate these information.
[0096] Compared with the prior art, the method specially configures a highlight reasoning engine based on iterative reinforcement learning for the multi-modal large model. The highlight reasoning engine can effectively guide the model to focus on key information, improve the performance of the model, and further improve the pertinence and efficiency of highlight extraction of the method.
[0097] For example, the spatio-temporal cross-attention mechanism model expression of the embodiment is as follows:
[0098]
[0099] In the formula, represents the attention score matrix; represents the historical fragment feature; represents the current fragment feature; Q t represents the query vector, which is generated based on the information at the current time step t, and is used to query in the historical and current fragment features; represents the splicing operation of the historical fragment feature and the current fragment feature to form a more comprehensive feature representation; d represents the latitude of the feature.
[0100] Through the design of the above-mentioned spatio-temporal cross-attention mechanism model, the method has the following advantages:
[0101] Firstly, it achieves cross-segment temporal alignment. By calculating attention weights, this mechanism can identify associated features in historical and current segments, thus achieving temporal alignment between different segments. This is crucial for understanding temporal dependencies and event development processes in multi-modal data.
[0102] Secondly, it can capture multi-modal associations. The spatio-temporal cross-attention mechanism can consider information in both time and space dimensions, effectively capturing complex associations between multi-modal data. For example, in video analysis, it can combine visual features, audio features, and text descriptions of video frames to more accurately locate and understand events in the video.
[0103] Thirdly, it can improve the accuracy of highlight point positioning. By focusing on key features and information, this mechanism helps improve the accuracy of multi-modal temporal positioning. The spatio-temporal cross-attention mechanism model can filter out irrelevant information, reduce noise interference, and make the multi-modal large model pay more attention to features useful for the positioning task.
[0104] Embodiment Three
[0105] As shown in Figure 2 Embodiment One or Two, this embodiment discloses a highlight point extraction system based on a multi-modal model, including a processor and a memory, the memory stores computer program instructions, when the computer program instructions are executed by the processor, the highlight point extraction method based on the multi-modal model recorded in Embodiment One or Embodiment Two is realized.
[0106] The system also includes a communication bus and a communication interface, as well as other components known to those skilled in the art, the settings and functions of which are known in the art, and therefore will not be described here. In the description of this specification, the meaning of "multiple" is at least two, for example, two, three or more, etc., unless otherwise explicitly specified.
[0107] Although the present specification has shown and described several embodiments of the present application, it will be apparent to those skilled in the art that many modifications, changes and substitutions can be made without departing from the idea and spirit of the present application. It should be understood that various alternatives to the embodiments of the present application described herein can be employed in practicing the present application.
Claims
1. A high light point extraction method based on a multi-modal model, characterized in that, The method comprises the following steps: in response to the input of video material, the video material is segmented into multiple overlapping sub-segments; inputting the multiple overlapping sub-segments into a preset multi-modal large model for joint feature extraction to obtain mixed features; extracting historical segment features and current segment features from the mixed features; inputting the historical segment features and the current segment features into a preset spatio-temporal cross-attention mechanism model to calculate an attention score matrix in the time dimension; taking the highlight segments or highlight nodes corresponding to the elements with scores greater than a preset threshold in the attention score matrix as highlight highlights; The specific expression of the mixed features obtained by inputting the multiple overlapping sub-segments into the preset multi-modal large model for joint feature extraction is: In the formula, represents the first mixed feature of the video material, represents a multimodal large model, represents the first overlapping sub-fragment of the video material, represents a cross-modal attention fusion operator, represents the first visual modality feature in the video material, represents the first audio modality feature in the video material, represents the first text or subtitle modality feature in the video material, Before inputting the multiple overlapping sub-segments into the preset multi-modal large model for joint feature extraction to obtain mixed features, the method further comprises: configuring a highlight inference engine, the highlight inference engine is used to drive the multi-modal large model to pay attention to the continuity of the plot of the video material; The highlight inference engine includes a Prompt dynamic evolution mechanism, and the specific expression of the Prompt dynamic evolution mechanism is: In the formula, represents the first iteration generated Prompt, represents the base Prompt, represents the splicing function, represents the summary feature of the previous reasoning, is used to summarize the key information of the multimodal data in the previous reasoning process; The highlight inference engine also sets a multi-modal reward function for defining double-objective optimization, wherein the defined double-objective optimization includes a continuity reward of adjacent overlapping sub-segments and an audience appeal reward; The definition of the multi-modal reward function is: wherein, denotes a maximization operation over parameters denotes an expectation operator, denotes a summation over time steps to denotes a discount factor, denotes a coherence reward for overlapping subsegments in state at time step denotes a hyperparameter, denotes an audience appeal reward for overlapping subsegments in state at time step 2. The high light point extraction method based on a multi-modal model according to claim 1, characterized in that, Before inputting the multiple overlapping sub-segments into the preset multi-modal large model for joint feature extraction to obtain mixed features, the method further comprises: defining a sub-segment overlap rate of the multiple overlapping sub-segments; According to the overlap rate, set a long-range dependence constraint for adjacent overlapping sub-segments.
3. The high light point extraction method based on a multi-modal model according to claim 2, characterized in that, The expression of the long-range dependence constraint is: In the formula, denotes the tile overlap ratio, denotes the first overlapping sub-tile of the video material, is temporally adjacent neighboring overlapping sub-tiles. 4.The high light point extraction method based on a multi-modal model according to claim 1, wherein, The spatio-temporal cross-attention mechanism model adopts any one of a decomposition attention framework, a supervised learning framework or a unified attention framework.
5. A high light point extraction system based on a multi-modal model, characterized by, The method comprises a processor and a memory, the memory stores computer program instructions, when the computer program instructions are executed by the processor, the method for extracting highlight highlights based on a multi-modal model is realized.
Citation Information
Patent Citations
Video content analysis method, equipment, medium and product based on deep learning
CN119763014A
Video timestamp event identification and reasoning method based on multi-modal large model
CN119723431A
Large model improved feature extraction-based wonderful lens detection method and system
CN119741638A