Video highlight detection method and device, medium, equipment and product

By combining multimodal feature fusion of video frames, audio, and text information and using an attention mechanism for modal alignment, the problem of inaccurate video highlight detection in existing technologies is solved, and more efficient highlight segment recognition is achieved.

CN121010929APending Publication Date: 2025-11-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129447.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies for video highlight detection are not accurate enough and fail to effectively utilize the multimodal information of videos, especially ignoring the temporal correspondence between textual and modal information.

Method used

By acquiring multimodal information from the video, including video frames, audio information, and text information, features are extracted using a text encoder and a visual encoder. Modality alignment is performed through an attention mechanism, and a mask matrix is ​​generated for feature fusion to ensure temporal correspondence.

Benefits of technology

It improves the accuracy of video highlight detection, enabling more accurate identification of exciting segments in videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010929A_ABST
    Figure CN121010929A_ABST
Patent Text Reader

Abstract

A video highlight detection method and apparatus, a medium, a device and a product, the video highlight detection method comprising: acquiring multi-modal information of a video, the multi-modal information comprising a video frame, audio information and text information, the text information comprising text content and time information of the text content in the video; determining video features and text features of the video according to the multi-modal information; according to the time information, performing modal alignment on the video features and the text features to obtain target video features after modal alignment; and determining a highlight segment of the video according to the target video feature. According to the technical scheme, the target video features can be obtained by fusing the video features and the text features with time relevance, the accuracy of fusion of the video features and the text features is ensured, and the accuracy of video highlight detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a video highlight detection method, apparatus, medium, device, and product. Background Technology

[0002] Video highlights refer to the most exciting or noteworthy segments in a video. Video highlight detection identifies these exciting or highlighting segments. For lengthy videos, users may not be interested in watching the entire video, while highlight segments are more likely to attract viewers and facilitate video sharing. However, the highlight segments detected by related technologies may not be accurate enough. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides a video highlight detection method, the method comprising: The video's multimodal information is acquired, including video frames, audio information, and text information, wherein the text information includes text content and the time information of the text content within the video. Based on the multimodal information, determine the video features and text features of the video; Based on the time information, modal alignment is performed on the video features and the text features to obtain modally aligned target video features; Based on the target video features, the highlight segments of the video are determined.

[0005] Secondly, this disclosure provides a video highlight detection device, the device comprising: The acquisition module is used to acquire multimodal information of the video, the multimodal information including video frames, audio information and text information, the text information including text content and the time information of the text content in the video; The feature determination module is used to determine the video features and text features of the video based on the multimodal information; The modal alignment module is used to perform modal alignment on the video features and the text features based on the time information to obtain modally aligned target video features; The segment determination module is used to determine the highlight segments of the video based on the characteristics of the target video.

[0006] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the video highlight detection method provided in the first aspect of this disclosure.

[0007] Fourthly, this disclosure provides an electronic device, comprising: A storage device on which computer programs are stored; A processing device is configured to execute the computer program in the storage device to implement the steps of the video highlight detection method provided in the first aspect of this disclosure.

[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the video highlight detection method provided in the first aspect of this disclosure.

[0009] The above technical solution utilizes multimodal information from video for highlight detection. This multimodal information includes video frames, audio information, and text information. The text information includes the text content and its temporal context within the video. The text content reflects key moments in the video, thus it is combined with the text information for highlight detection. Furthermore, based on the temporal information, video and text features are modally aligned to obtain modally aligned target video features. This ensures that the target video features are obtained by fusing video features with temporally correlated text features, guaranteeing the accuracy of the feature fusion and improving the accuracy of highlight detection.

[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a video highlight detection method according to an exemplary embodiment.

[0012] Figure 2 This is a schematic diagram illustrating a video highlight detection method according to an exemplary embodiment.

[0013] Figure 3 This is a schematic diagram of a mask matrix as an example.

[0014] Figure 4This is a flowchart illustrating a method for determining highlight segments of a video based on target video features, according to an exemplary embodiment.

[0015] Figure 5 This is a block diagram illustrating a video highlight detection device according to an exemplary embodiment.

[0016] Figure 6 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0019] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0025] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0027] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0028] Figure 1 This is a flowchart illustrating a video highlight detection method according to an exemplary embodiment. The video highlight detection method can be applied to electronic devices, such as terminal devices or server devices. Figure 1 As shown, the video highlight detection method includes steps 11 to 14.

[0029] Step 11: Obtain multimodal information from the video.

[0030] The video can be any short or long video, published or unpublished, without any restrictions. Multimodal information includes video frames, audio information, and text information. The text information includes the text content and the timing of the text content within the video.

[0031] The video can be segmented at a preset FPS (Frames Per Second) to obtain multiple video frames. To achieve precise highlight positioning, the segmentation frequency can be increased. For example, if the video duration is 10 seconds, 20 video frames can be obtained by segmenting at FPS=2.

[0032] Audio extraction from a video yields audio information, which can include multiple audio segments. For example, if an audio segment is 1 second long and the video is 10 seconds long, the audio information will consist of 10 audio segments.

[0033] If the video contains subtitles, the text content can be the subtitle text, obtained by performing Optical Character Recognition (OCR) on the subtitle text in the video frames. If the video does not contain subtitles, the text content can be obtained through speech recognition, i.e., by performing Automatic Speech Recognition (ASR) on the extracted audio information. The text content can be a single character, a word, or a sentence. For example, if the video contains the subtitle "What is this?", then the sentence "What is this?" is considered text content. The timing information of the text content in the video can include the start and end times. If the text content is obtained from the subtitle text, the start time is the time when the subtitle text begins to appear, and the end time is the time when the subtitle text begins to disappear. If the text content is obtained through speech recognition, the start time is the time when the text content begins to be spoken in the video, and the end time is the time when the corresponding audio in the video ends.

[0034] Step 12: Determine the video features and text features of the video based on the multimodal information.

[0035] Video features can be obtained from video frames and audio information, while text features can be obtained by extracting features from text content.

[0036] Step 13: Based on the time information, perform modal alignment on the video features and text features to obtain the modally aligned target video features.

[0037] Step 14: Determine the highlight segments of the video based on the characteristics of the target video.

[0038] When performing video highlight recognition in related technologies, on the one hand, only the visual image information and audio information of the video are usually used, without considering the text information of the video. However, text information has high information value and can reflect the highlights of the video. On the other hand, the temporal correspondence between different modal information is usually ignored.

[0039] In this disclosure, specular highlight detection is performed using multimodal information from video, including video frames, audio information, and text information. Furthermore, the temporal information can be a time interval consisting of a start time and an end time. Based on this temporal information, modal alignment is performed on video features and text features to ensure the temporal correspondence of feature fusion. Modal alignment can be achieved, for example, using an attention mechanism. The video features corresponding to the text content within the time interval in the video are temporally correlated with the text features of that text content. Modal alignment allows the corresponding video features within the time interval to mutually influence the text features, while the text features are not temporally correlated with video features corresponding to other times, thus avoiding the text features from being correlated with video features corresponding to other times.

[0040] After modal alignment of video and text features, fused target video and text features can be obtained. When performing highlight detection, the target video features can be used to detect highlight segments in the video.

[0041] The above technical solution utilizes multimodal information from video for highlight detection. This multimodal information includes video frames, audio information, and text information. The text information includes the text content and its temporal context within the video. The text content reflects key moments in the video, thus it is combined with the text information for highlight detection. Furthermore, based on the temporal information, video and text features are modally aligned to obtain modally aligned target video features. This ensures that the target video features are obtained by fusing video features with temporally correlated text features, guaranteeing the accuracy of the feature fusion and improving the accuracy of highlight detection.

[0042] Figure 2 This is a schematic diagram illustrating a video highlight detection method according to an exemplary embodiment. (Refer to...) Figure 2 This paper details the video highlight detection method disclosed herein.

[0043] Step 12, based on multimodal information, determines the video features and text features of the video, which may include: Determine video features based on video frames and audio information; Feature extraction is performed on the text content to obtain text features.

[0044] like Figure 2As shown, feature extraction from text content can be achieved by inputting the text content into a text encoder to obtain text features (text embeddings) output by the encoder. For example, if the text content consists of M subtitles (i.e., the video contains M subtitles, each of which can be a sentence), then one text content corresponds to one text feature, resulting in M ​​text features. The dimension of the text features output by the text encoder is, for example, f4. Inputting these f4-dimensional text features into a multilayer perceptron (MLP) allows the MLP to project the dimensions of the text features onto a dimension f3. Dimension f3 can be a preset dimension; for example, in a modality alignment module including an attention module, f3 can be the dimension of the hidden state of the attention module, which is also known as `hidden_size`. This projects the dimensions of the text features onto the dimensions required by the attention module, allowing the attention module to process the text features.

[0045] The text encoders mentioned above can be Roberta-zh (Robustly Optimized BERT Approach for Chinese), Bert (Bidirectional Encoder Representations from Transformers), or LLM (Large Language Model).

[0046] The following describes an implementation method for determining video features based on video frames and audio information.

[0047] In the first embodiment, visual features can be obtained by extracting features from video frames using an image encoder, and audio features can be obtained by extracting features from audio information using an audio encoder. The visual and audio features can then be fused, for example, using a cross-attention mechanism, to obtain video features.

[0048] In the second embodiment, determining video features based on video frames and audio information may include: if the number of video frames is inconsistent with the number of audio segments in the audio information, then performing an interpolation operation using video frames or audio segments to make the number of audio segments consistent with the number of video frames after the interpolation operation; adding the visual features of the video frames with the audio features of the corresponding audio segments to obtain the video features.

[0049] For example, if the number of audio segments is less than the number of video frames, an interpolation operation is performed on the audio segments to make the number of audio segments consistent with the number of video frames after the interpolation operation.

[0050] like Figure 2 As shown, the audio information contains k audio segments and N video frames, where k < N. Audio segment interpolation can be used to obtain N audio segments, some of which may be repeated. For example, if the audio segment duration is 1 second and the video duration is 10 seconds, the audio information includes 10 audio segments, and the video frames are 20. Audio segment interpolation can be used to make the total number of audio segments 20. For instance, these 20 audio segments correspond to 20 positions. The interpolation operation could be to copy the first audio segment and insert it into the second position (between the original first and second audio segments), insert the second audio segment into the fourth position (between the original second and third audio segments), and so on.

[0051] If the number of audio segments is the same as the number of video frames, then no further interpolation is needed.

[0052] For example, if the number of video frames is less than the number of audio segments, video frame interpolation is used to ensure that the number of interpolated video frames matches the number of audio segments. For instance, if frame extraction is performed at FPS=0.5 (one video frame every 2 seconds), a 10-second video will have 5 video frames. Since there are 10 audio segments, video frame interpolation can be used to obtain 10 video frames, some of which may be duplicates. For example, if the 10 video frames correspond to 10 positions, the interpolation operation could be to insert the first video frame into the second position (between the original first and second video frames), insert the second video frame into the fourth position (between the original second and third video frames), and so on. Subsequent operations follow the same principle.

[0053] like Figure 2 As shown, video frames are input to a visual encoder to obtain visual features output by the visual encoder. N video frames correspond to N visual features, and the dimension of the visual features output by the visual encoder is f1 (e.g., 512 dimensions). Audio segments are input to an audio encoder to obtain audio features output by the audio encoder. N audio segments correspond to N audio features, and the dimension of the audio features output by the audio encoder is f2 (e.g., 128 dimensions). To align the dimensions of the visual features with the dimensions of the audio features, a multilayer perceptron can be used to project the audio features of dimension f2 onto dimension f1 to obtain audio features of dimension f1, meaning that each audio feature has a dimension of f1.

[0054] Next, visual features of the same dimension are added to their corresponding audio features to obtain video features of dimension f1. This feature addition can refer to adding the first visual feature to the first audio feature to obtain the first video feature. The same process is applied to other features, resulting in N video features, each with dimension f1. Then, using an MLP, the video features of dimension f1 are projected onto dimension f3 to obtain N video features of dimension f3. Dimension f3, as explained above, represents the dimension of the hidden state of the attention module.

[0055] In the above embodiment, the dimension f2 is less than f1, that is, the audio feature with dimension f2 needs to be transformed by MLP. If the dimension f1 is less than the dimension f2, the visual feature with dimension f1 can be transformed by MLP to project the dimension of the visual feature to dimension f2.

[0056] Thus, by aligning the number of video frames with the number of audio segments, and then adding visual and audio features of the same dimension, the visual and audio features are fused to obtain video features. These video features integrate both visual and audio characteristics of the video. In this disclosure, experiments have shown that using the second embodiment—feature addition—to fuse visual and audio features results in higher accuracy for highlight detection. Therefore, this disclosure can employ the second embodiment to determine video features based on video frames and audio information.

[0057] The visual encoder mentioned above can be one of the following: CLIP (Contrastive Language-Image Pre-Training), ResNet (Residual Network) based on CNN (Convolutional Neural Networks), Inception-v1 based on CNN, or Swin (Shifted Window Transformer) based on ViT (Vision Transformer). The audio encoder can be one of the following: VGGish (Visual Geometry Group style Audio Classification Model), PANNS (Pretrained Audio Neural Network), or Wav2vec.

[0058] The following describes an implementation method for modal alignment of video features and text features in this disclosure.

[0059] Step 13 may include: generating a mask matrix based on time information; and performing modal alignment of video features and text features based on the mask matrix to obtain target video features.

[0060] The number of rows and columns of the mask matrix is ​​the sum of the number of video features and the number of text features, that is, the size of the mask matrix is ​​(N+M)×(N+M). Figure 3 This is a schematic diagram of an exemplary mask matrix, such as... Figure 3 As shown, the number of video features is 6 and the number of text features is 3. The size of the mask matrix is ​​9×9, where video1 to video6 represent video features and text1 to text3 represent text features.

[0061] The value of the target element in the mask matrix is ​​the first value, which can be 0. The values ​​of the elements in the mask matrix other than the target element are the third values, which can be negative infinity (represented by -inf).

[0062] The target elements include the following elements: elements whose behavior is the first video feature and whose column is the first text feature, elements whose behavior is the first text feature and whose column is the first video feature, and the time corresponding to the first video feature belongs to the time interval formed by the start time and end time corresponding to the first text feature.

[0063] For example, taking "text1" as the first text feature, "text1" is obtained by extracting features from text content 1. The time information of text content 1 in the video is [0s, 1s], that is, the start time corresponding to "text1" is 0s, and the end time corresponding to "text1" is 1s. Each video frame has a corresponding time. For example, video frame 1 and video frame 2 are obtained by extracting frames from the [0s, 1s] time interval of the video. "video1" is obtained based on the visual features of video frame 1, and "video2" is obtained based on the visual features of video frame 2. Therefore, the time corresponding to the first video features "video1" and "video2" belongs to the time interval [0s, 1s]. Figure 3 As shown, the elements in the first video feature and in the first text feature are located in the same row, including the element in row 1, column 7, and the element in row 2, column 7. The elements in the first text feature and in the first video feature are located in the same row, including the element in row 7, column 1, and the element in row 7, column 2.

[0064] Using text2 as the first text feature can be understood as explained above. For example... Figure 3As shown, the elements in the row that are the first video feature and the elements in the column that are the first text feature include the elements in the 3rd row and 8th column, and the elements in the 4th row and 8th column. The elements in the row that are the first text feature and the elements in the column that are the first video feature include the elements in the 8th row and 3rd column, and the elements in the 8th row and 4th column.

[0065] Using text3 as the first text feature can be understood in accordance with the above explanation, such as... Figure 3 As shown, the elements in the row that are the first video feature and the elements in the column that are the first text feature include the elements in the 5th row and 9th column, and the elements in the 6th row and 9th column. The elements in the row that are the first text feature and the elements in the column that are the first video feature include the elements in the 9th row and 5th column, and the elements in the 9th row and 6th column.

[0066] In this way, during subsequent feature fusion based on the mask matrix, only features from different modalities within the same time interval can mutually focus on each other, thus ensuring the temporal correspondence between video and text features during feature fusion. That is, for example, subtitle text appearing at [0s, 1s] is temporally related to video frames within [0s, 1s], while subtitle text appearing at [0s, 1s] is not related to video frames at other times, such as [9s, 10s]. Therefore, the mask matrix allows for the fusion of temporally related video and text features.

[0067] Additionally, target elements may include the following: elements whose row and column are both video features, and elements whose row and column are both text features. During attention alignment, features of the same modality can also mutually focus, such as... Figure 3 As shown, elements whose rows and columns are both video features include elements with rows (video1 to video6) and columns (video1 to video6), while elements whose rows and columns are both text features include elements with rows (text1 to text3) and columns (text1 to text3). This allows for the fusion of features of the same modality when aligning video and text feature modalities.

[0068] In this disclosure, the modality alignment module can adopt a Transformer structure. Transformer is a neural network architecture based on a self-attention mechanism, using Transformer autocorrelation calculation for modality alignment. The Transformer structure can contain multiple layers, each layer including a multi-head self-attention module (MHA), a feed-forward network (FFN), residual connections, and a layer normalization module. The FFN can include FFNs for processing text features and FFNs for processing video features. Information transfer between multiple layers can refer to the Transformer calculation methods in related technologies.

[0069] The process involves concatenating video and text features, resulting in N+M concatenated features. These are then processed by the Embedding layer to obtain Segment Embedding, Position Embedding, and Feature Embedding. Feature Embedding can include text-related and video-related features. Text-related features may include CLS markers, each text feature (e.g., represented by T0, T1, etc.), and SEP markers. Video-related features may include each video feature (e.g., represented by V0, V1, etc.) and PAD markers. CLS markers are typically added to the beginning of the text sequence, SEP markers are used to segment the text, and PAD markers serve as padding and placeholders. Position Embedding provides the positional information of each element in the sequence. Segment Embedding distinguishes different features by assigning a unique embedding vector to each feature. The Embedding obtained after processing by the Embedding layer is then input into the modality alignment module.

[0070] The attention module represents each feature as a query (Q), key (K), and value (V), and calculates attention scores between features to form an attention score matrix. The size of the attention score matrix is ​​the same as the size of the mask matrix. Attention scores measure the relevance between one feature and another. Typically, the attention score is calculated by the dot product of the query of one feature and the key of another. The more similar the query of one feature is to the key of another, the higher the corresponding attention score, indicating that these two features require more attention.

[0071] Among them, modal alignment is performed on video features and text features based on the mask matrix to obtain target video features, which may include: Based on the mask matrix, an attention weight matrix is ​​determined, wherein the size of the attention weight matrix is ​​the same as that of the mask matrix. The value of the element corresponding to the target element in the attention weight matrix is ​​a second value, which is determined based on the correlation between the features of the row where the target element is located and the features of the column where the target element is located. Based on the attention weight matrix, the video features and text features are fused to obtain the target video features.

[0072] Specifically, the attention score matrix is ​​added to the mask matrix, meaning corresponding elements are added together. The resulting matrix is ​​then processed using softmax to obtain the attention weight matrix. The values ​​of other elements in the attention weight matrix are close to 0; these other elements include all elements except those corresponding to the target element.

[0073] like Figure 3 As shown, for example, the element in the 1st row and 7th column of the mask matrix is ​​0, and the value of the element in the 1st row and 7th column of the attention score matrix is ​​obtained based on the correlation between feature video1 and feature text1, that is, obtained in the way the attention score is calculated as described above. After adding the attention score matrix and the mask matrix, the element in the 1st row and 7th column is the attention score.

[0074] For example, the element in the 1st row and 8th column of the mask matrix is ​​-inf. The value of the element in the 1st row and 8th column of the attention score matrix is ​​obtained based on the correlation between feature video1 and feature text2. After adding the attention score matrix and the mask matrix, the element in the 1st row and 8th column is -inf. After softmax processing, the element in the 1st row and 8th column of the attention weight matrix approaches 0. In this way, there is no temporal correlation between feature video1 and feature text2. Therefore, when performing modal alignment, since the corresponding attention weights are close to 0, the two will not pay attention to each other.

[0075] The method for fusing video features and text features based on the attention weight matrix can be to calculate the value (V) of the features by weighting them according to the attention weight matrix. For details, please refer to the fusion calculation of attention mechanism in related technologies.

[0076] After multiple layers of computation in the modal alignment module, the modal alignment module can output fused features. The number of fused features is N+M, including N fused target video features and M target text features. The N target video features can be used for specular prediction.

[0077] The above technical solution, which uses the generated mask matrix for feature fusion and employs an attention mechanism for feature fusion, enables features of the same modality to pay attention to each other, and ensures that only features of different modalities within the same time interval can pay attention to each other. This guarantees the temporal correspondence between video features and text features during feature fusion, and improves the accuracy of specular prediction based on the fused target video features.

[0078] The following describes the implementation method of determining highlight segments based on target video features in this disclosure, which can be derived from... Figure 2 The highlight prediction module shown identifies highlight segments. Figure 4 This is a flowchart illustrating a method for determining highlight segments of a video based on target video features, according to an exemplary embodiment. Figure 4 As shown, step 14 includes steps 141 to 144.

[0079] Step 141: Determine the first highlight score for each video frame based on the target video features.

[0080] For example, N target video features can be input into a pre-trained specular prediction model. The specular prediction model can be in the form of an MLP. The specular prediction model can output the first specular score of each video frame. The first specular score can represent the level of excitement of the video frame. The higher the first specular score, the higher the level of excitement of the video frame.

[0081] Step 142: Perform scene segmentation on the video to obtain multiple video clips.

[0082] For example, the segmentation process can employ a uniform segmentation method or a segmentation method based on changes in image scenes or features. For instance, if the video duration is 20 seconds, multiple video segments can include video segments of [0s, 5s], [6s, 10s], [11s, 15s], and [16s, 20s].

[0083] Step 143: Determine the first video segment from multiple video segments.

[0084] The first video segment includes a first video frame, and the first highlight score corresponding to the first video frame satisfies a first preset condition.

[0085] The first preset condition can be met if the first highlight score is ranked in descending order and is in the first m1 positions, or if the first preset condition is met if the first highlight score is greater than the first score threshold. Interpreting this in terms of video frames corresponding to time, the first video frame includes video frames at 1s, 2s, 3s, 7s, and 9s. Therefore, the first video segment includes video segments at [0s, 5s] and [6s, 10s].

[0086] Step 144: Determine the highlight segment based on the first video segment.

[0087] The implementation method for step 144 can be as follows: Determine the second highlight score for each first video segment, wherein, when the first video segment includes multiple first video frames, the second highlight score is the average of the first highlight scores corresponding to the multiple first video frames respectively, and when the first video segment includes one first video frame, the second highlight score is the first highlight score of the first video frame.

[0088] For example, the second highlight score of a video clip [0s, 5s] is the average of the first highlight scores corresponding to the video frames at 1s, 2s, and 3s, respectively. The second highlight score of a video clip [6s, 10s] is the average of the first highlight scores corresponding to the video frames at 7s and 9s, respectively. For instance, if a video clip [6s, 10s] contains a first video frame, such as the 7th video frame, then the second highlight score of the video clip [6s, 10s] is the first highlight score of the 7th video frame.

[0089] A highlight segment is obtained from the first video segment whose second highlight score meets the second preset condition.

[0090] The second preset condition can be met if the second highlight score is ranked in the first m2 positions in descending order, or if the second highlight score is greater than the second score threshold. m1 and m2 are positive integers and can be the same or different. The pre-defined first score threshold and second score threshold can be the same or different.

[0091] In one embodiment, if there are multiple first video segments that meet the second preset condition, the multiple first video segments are spliced ​​together in chronological order to obtain a highlight segment.

[0092] In one embodiment, the start and end times of the first video segment that meets the second preset condition are updated. The start time is updated to the minimum time of the first video frame contained in the first video segment, and the end time is updated to the maximum time of the first video frame contained in the first video segment. This further filters out the portion of the first video segment that does not contain highlight video frames. For example, a video segment of [0s, 5s] is updated to a video segment of [1s, 3s], and a video segment of [6s, 10s] is updated to a video segment of [7s, 9s]. For example, if there are multiple first video segments that meet the second preset condition, the multiple updated first video segments are spliced ​​together in chronological order to obtain a highlight segment. Alternatively, the segment between the minimum and maximum times of the multiple updated first video segments can be used as a highlight segment. Following the above example, a video segment of [1s, 9s] can be used as a highlight segment.

[0093] When the video highlight detection method is applied to terminal devices, after identifying the highlight segments, the highlight segments can be played by the terminal devices for user viewing. When the video highlight detection method is applied to server devices, after identifying the highlight segments, the server can send the highlight segments to the terminal devices, which can then play them for user viewing.

[0094] The above technical solution utilizes multimodal information from video for video highlight detection. Furthermore, based on temporal information, modal alignment is performed on video features and text features to ensure the accuracy of video feature fusion and improve the accuracy of video highlight detection.

[0095] Based on the same inventive concept, this disclosure also provides a video highlight detection device. Figure 5 This is a block diagram illustrating a video highlight detection apparatus according to an exemplary embodiment, such as... Figure 5 As shown, the video highlight detection device 50 may include: The acquisition module 51 is used to acquire multimodal information of the video, the multimodal information including video frames, audio information and text information, the text information including text content and the time information of the text content in the video; Feature determination module 52 is used to determine the video features and text features of the video based on the multimodal information; Modality alignment module 53 is used to perform modality alignment on the video features and the text features based on the time information to obtain modally aligned target video features; The segment determination module 54 is used to determine the highlight segments of the video based on the characteristics of the target video.

[0096] Optionally, the time information includes start time and end time; the modal alignment module 53 includes: A generation submodule is used to generate a mask matrix based on the time information, wherein the number of rows and columns of the mask matrix is ​​the sum of the number of video features and the number of text features, the value of the target element in the mask matrix is ​​a first value, and the target element includes the following elements: the element whose row is the first video feature and whose column is the first text feature, the element whose row is the first text feature and whose column is the first video feature, and the time corresponding to the first video feature belongs to the time interval formed by the start time and the end time corresponding to the first text feature; The alignment submodule is used to perform modal alignment of the video features and the text features according to the mask matrix to obtain the target video features.

[0097] Optionally, the target element may further include the following elements: elements whose row and column are both video features, and elements whose row and column are both text features.

[0098] Optionally, the alignment submodule includes: The first determining submodule is used to determine an attention weight matrix based on the mask matrix, wherein the size of the attention weight matrix is ​​the same as the size of the mask matrix, and the value of the element corresponding to the target element in the attention weight matrix is ​​a second value, which is determined based on the correlation between the features of the row where the target element is located and the features of the column where the target element is located. The fusion submodule is used to fuse the video features and the text features according to the attention weight matrix to obtain the target video features.

[0099] Optionally, the feature determination module 52 includes: The second determining submodule is used to determine the video features based on the video frames and the audio information; The extraction submodule is used to extract features from the text content to obtain the text features.

[0100] Optionally, the second determining submodule includes: The interpolation operation submodule is used to perform an interpolation operation using either the video frames or the audio segments if the number of video frames is inconsistent with the number of audio segments in the audio information, so that the number of audio segments after the interpolation operation is consistent with the number of video frames. The feature addition submodule is used to add the visual features of the video frame to the audio features of the corresponding audio segment to obtain the video features.

[0101] Optionally, the fragment determination module 54 includes: The third determining submodule is used to determine the first highlight score of each video frame based on the target video features; The submodule for segmentation processing is used to process the video into multiple video segments. The fourth determining submodule is used to determine a first video segment from the plurality of video segments, the first video segment including a first video frame, and the first highlight score corresponding to the first video frame satisfies a first preset condition; The fifth determining submodule is used to determine the highlight segment based on the first video segment.

[0102] Optionally, the fifth determining submodule includes: The sixth determining submodule is used to determine the second highlight score of each of the first video segments, wherein, when the first video segment includes multiple first video frames, the second highlight score is the average of the first highlight scores corresponding to the multiple first video frames respectively, and when the first video segment includes one first video frame, the second highlight score is the first highlight score of the first video frame. The seventh determining submodule is used to obtain the highlight segment based on the first video segment whose second highlight score meets the second preset condition.

[0103] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0104] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0105] like Figure 6As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0106] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0107] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0108] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0109] In some implementations, terminals and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0110] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0111] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire multimodal information of a video, the multimodal information including video frames, audio information, and text information, the text information including text content and time information of the text content in the video; Based on the multimodal information, determine the video features and text features of the video; Based on the time information, modal alignment is performed on the video features and the text features to obtain modally aligned target video features; Based on the target video features, the highlight segments of the video are determined.

[0112] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as a "module for acquiring multimodal information".

[0115] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0116] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0117] According to one or more embodiments of this disclosure, Example 1 provides a video highlight detection method, the method comprising: The video's multimodal information is acquired, including video frames, audio information, and text information, wherein the text information includes text content and the time information of the text content within the video. Based on the multimodal information, determine the video features and text features of the video; Based on the time information, modal alignment is performed on the video features and the text features to obtain modally aligned target video features; Based on the target video features, the highlight segments of the video are determined.

[0118] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the time information includes a start time and an end time; the step of performing modal alignment on the video features and the text features based on the time information to obtain modally aligned target video features includes: Based on the time information, a mask matrix is ​​generated, wherein the number of rows and columns of the mask matrix is ​​the sum of the number of video features and the number of text features. The value of the target element in the mask matrix is ​​a first value. The target element includes the following elements: the element whose row is the first video feature and whose column is the first text feature, and the element whose row is the first text feature and whose column is the first video feature. The time corresponding to the first video feature belongs to the time interval formed by the start time and the end time corresponding to the first text feature. Based on the mask matrix, modal alignment is performed on the video features and the text features to obtain the target video features.

[0119] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein the target element further includes the following elements: elements whose row and column are both the video features, and elements whose row and column are both the text features.

[0120] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein modal alignment of the video features and the text features based on the mask matrix to obtain the target video features includes: Based on the mask matrix, an attention weight matrix is ​​determined, wherein the size of the attention weight matrix is ​​the same as the size of the mask matrix, and the value of the element corresponding to the target element in the attention weight matrix is ​​a second value, which is determined based on the correlation between the features of the row where the target element is located and the features of the column where the target element is located. The video features and the text features are fused according to the attention weight matrix to obtain the target video features.

[0121] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein determining the video features and text features of the video based on the multimodal information includes: The video features are determined based on the video frames and the audio information; The text content is subjected to feature extraction to obtain the text features.

[0122] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein determining the video features based on the video frame and the audio information includes: If the number of video frames is inconsistent with the number of audio segments in the audio information, then an interpolation operation is performed using either the video frames or the audio segments to make the number of audio segments consistent with the number of video frames after the interpolation operation. The video features are obtained by adding the visual features of the video frame to the audio features of the corresponding audio segment.

[0123] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein determining the highlight segment of the video based on the target video features includes: Based on the target video features, determine the first highlight score for each video frame; The video is split into multiple video segments; A first video segment is determined from the plurality of video segments, the first video segment includes a first video frame, and the first highlight score corresponding to the first video frame satisfies a first preset condition; The highlight segment is determined based on the first video segment.

[0124] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 7, wherein determining the highlight segment based on the first video segment includes: A second highlight score is determined for each of the first video segments, wherein, when the first video segment includes multiple first video frames, the second highlight score is the average of the first highlight scores corresponding to the multiple first video frames respectively, and when the first video segment includes one first video frame, the second highlight score is the first highlight score of the first video frame. The highlight segment is obtained from the first video segment whose second highlight score satisfies the second preset condition.

[0125] According to one or more embodiments of this disclosure, Example 9 provides a video highlight detection apparatus, the apparatus comprising: The acquisition module is used to acquire multimodal information of the video, the multimodal information including video frames, audio information and text information, the text information including text content and the time information of the text content in the video; The feature determination module is used to determine the video features and text features of the video based on the multimodal information; The modal alignment module is used to perform modal alignment on the video features and the text features based on the time information to obtain modally aligned target video features; The segment determination module is used to determine the highlight segments of the video based on the characteristics of the target video.

[0126] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1 to 8.

[0127] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, including: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of any one of Examples 1 to 8.

[0128] According to one or more embodiments of the present disclosure, Example 12 provides a computer program product including a computer program that, when executed by a processor, implements the steps of the method described in any one of Examples 1 to 8.

[0129] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0130] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0131] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method of video highlight detection, the method comprising: The method comprises: acquiring multi-modal information of a video, the multi-modal information comprising video frames, audio information and text information, the text information comprising text content and time information of the text content in the video; determining video features and text features of the video according to the multi-modal information; aligning the video features and the text features according to the time information to obtain target video features after modal alignment; determining a highlight segment of the video according to the target video features.

2. The method of claim 1, wherein, The time information comprises a start time and an end time; and the aligning the video features and the text features according to the time information to obtain target video features after modal alignment comprises: generating a mask matrix according to the time information, wherein the number of rows and the number of columns of the mask matrix are the sum of the number of the video features and the number of the text features, and the value of a target element in the mask matrix is a first value, the target element comprising an element whose row is a first video feature and whose column is a first text feature, or an element whose row is the first text feature and whose column is the first video feature, the time corresponding to the first video feature belonging to a time interval constituted by the start time and the end time corresponding to the first text feature; aligning the video features and the text features according to the mask matrix to obtain the target video features.

3. The method of claim 2, wherein, The target element further comprises an element whose row and column are both the video features, or an element whose row and column are both the text features.

4. The method of claim 2, wherein, The aligning the video features and the text features according to the mask matrix to obtain the target video features comprises: determining an attention weight matrix according to the mask matrix, wherein the size of the attention weight matrix is the same as the size of the mask matrix, and the value of an element corresponding to the target element in the attention weight matrix is a second value, the second value being determined according to the correlation between the feature of the row where the target element is located and the feature of the column where the target element is located; fusing the video features and the text features according to the attention weight matrix to obtain the target video features.

5. The method of claim 1, wherein, The determining the video features and the text features of the video according to the multi-modal information comprises: determining the video features according to the video frames and the audio information; extracting features from the text content to obtain the text features.

6. The method of claim 5, wherein, The determining the video features according to the video frames and the audio information comprises: if the number of the video frames is inconsistent with the number of audio segments in the audio information, performing an interpolation operation on the video frames or the audio segments to make the number of the audio segments consistent with the number of the video frames after the interpolation operation; adding the visual features of the video frames and the audio features of the corresponding audio segments to obtain the video features.

7. The method of claim 1, wherein, The determining the highlight segment of the video according to the target video features comprises: According to the target video feature, a first highlight score of each video frame is determined; The video is shot by a script, and a plurality of video clips are obtained; A first video clip is determined from the plurality of video clips, the first video clip including a first video frame, and a first highlight score corresponding to the first video frame satisfying a first preset condition; According to the first video clip, the highlight clip is determined.

8. The method of claim 7, wherein, The determination of the highlight clip according to the first video clip includes: A second highlight score of each first video clip is determined, wherein, in the case that the first video clip includes a plurality of first video frames, the second highlight score is the average of the first highlight scores corresponding to the plurality of first video frames, and in the case that the first video clip includes one first video frame, the second highlight score is the first highlight score of the first video frame; According to the first video clip in which the second highlight score satisfies a second preset condition, the highlight clip is obtained.

9. A video highlight detection apparatus characterized by comprising: The device includes: An acquisition module configured to acquire multi-modal information of a video, the multi-modal information including video frames, audio information, and text information, the text information including text content and time information of the text content in the video; A feature determination module configured to determine video features and text features of the video according to the multi-modal information; A modal alignment module configured to perform modal alignment on the video features and the text features according to the time information, and obtain target video features after modal alignment; A clip determination module configured to determine a highlight clip of the video according to the target video features.

10. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by a processing device to implement the steps of the method in any one of claims 1-8.

11. An electronic device, comprising: It includes: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the method in any one of claims 1-8.

12. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the method in any one of claims 1-8.