Highlight Video Recognition Method and Device, Electronic Device, and Storage Medium
By encoding and splicing the visual and audio features of the video, and using the self-attention mechanism to establish audio and video feature correlation, the low accuracy problem caused by the asynchrony of audio and video in traditional highlight detection technology is solved, and the accuracy of recognition of highlight video clips is improved.
Patent Information
- Application Number
- CN202210615599.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-31
AI Technical Summary
Traditional high-light detection technology has low accuracy due to the insync of audio and video, and cannot effectively identify the highlight fragments in the video.
By extracting the visual and audio features of the video, encoded and splicing them in the head and tail, and using the self-attention mechanism to perform feature encoding, establish the correlation between audio and video features and identify highlight video clips.
It improves the recognition accuracy of highlight video clips, alleviates the problem of domain asymmetry between audio and video features, and enhances the ability to fusion information of different modalities.
Smart Images

Figure CN115035441B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a method and apparatus for high-light video recognition, an electronic device, and a storage medium. Background Art
[0002] Video high-light detection technology is mainly applied to automatic video editing, recommendation, retrieval, and various downstream application scenarios. By means of computer vision technology, exciting segments or highlight moments that have appeared in the video are located, and these segments / moments can effectively arouse the interest of viewers. For example, for a video of a basketball game, the algorithm outputs exciting segments such as exciting basketball goal segments and game-winning shot segments by jointly modeling video input and audio input.
[0003] Traditional high-light detection technology is based on the assumption of audio-visual synchronization. However, since the video picture and sound may not be synchronized in terms of high-light time, its accuracy needs to be further improved. Summary of the Invention
[0004] The present disclosure proposes a technical solution for high-light video recognition.
[0005] According to one aspect of the present disclosure, there is provided a method for high-light video recognition, including:
[0006] Extracting visual features and audio features of a video to be recognized, where the video to be recognized is segmented into multiple video segments, the visual features include visual sub-features of the multiple video segments arranged in time sequence, and the audio features include audio sub-features of the multiple video segments arranged in time sequence;
[0007] Encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features;
[0008] Performing head-to-tail splicing on the visual encoded features and audio encoded features to obtain spliced features;
[0009] Performing feature encoding on the spliced features based on a self-attention mechanism to obtain encoded spliced features;
[0010] Identifying high-light video segments among the multiple video segments based on the encoded spliced features.
[0011] In a possible implementation manner, encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features includes:
[0012] Extracting first global context features of each visual sub-feature in the visual features;
[0013] Fuse each of the first global context features with the corresponding visual sub-features to obtain a plurality of first visual sub-features as the visual encoding features;
[0014] Extract the second global context features of each audio sub-feature in the audio feature;
[0015] Fuse each of the second global context features with the corresponding audio sub-features to obtain a plurality of first audio sub-features as the audio encoding features.
[0016] In one possible implementation, the feature encoding of the concatenated feature based on the self-attention mechanism to obtain the encoded concatenated feature includes:
[0017] Extract the third global context features of each concatenated sub-feature in the concatenated feature, where the concatenated sub-feature is a first visual sub-feature or a first audio sub-feature;
[0018] Fuse each of the third global context features with the corresponding concatenated sub-features to obtain the encoded concatenated feature.
[0019] In one possible implementation, the identification of the highlight video segments among the multiple video segments based on the encoded concatenated feature includes:
[0020] At the concatenation position, split the encoded concatenated feature to obtain a second visual sub-feature and a second audio sub-feature;
[0021] Fuse the second visual sub-feature and the second audio sub-feature corresponding to the same video segment to obtain a plurality of fused sub-features;
[0022] Based on the fused sub-features, determine whether the segment corresponding to the fused sub-feature is a highlight video segment.
[0023] In one possible implementation, after encoding the visual feature and the audio feature respectively to obtain the visual encoding feature and the audio encoding feature, the method further includes:
[0024] Fuse the global highlight features extracted from the visual feature and the audio feature with the corresponding visual encoding feature and audio encoding feature respectively to obtain a visual fusion feature and an audio fusion feature;
[0025] The identification of the highlight video segments among the multiple video segments based on the encoded concatenated feature includes:
[0026] Identify the highlight video segments among the multiple video segments based on the encoded concatenated feature, the visual fusion feature, and the audio fusion feature.
[0027] In a possible implementation, the method for extracting global highlight features from the visual features and audio features includes:
[0028] Based on the cross-attention mechanism, using the global highlight embedding, extract the global highlight features in the visual encoded features and audio encoded features respectively. The global highlight embedding is a vector obtained through training for global abstraction and generalization of highlight features.
[0029] In a possible implementation, the method for identifying highlight video segments among the multiple video segments based on the encoded concatenated features, the visual fusion features, and the audio fusion features includes:
[0030] Based on the encoded concatenated features, obtain a first recognition result;
[0031] Based on the visual fusion features, obtain a second recognition result;
[0032] Based on the audio fusion features, obtain a third recognition result;
[0033] Perform weighted fusion on the first recognition result, the second recognition result, and the third recognition result to obtain the recognition result of the highlight segment.
[0034] In a possible implementation, the method for extracting the visual features and audio features of the video to be recognized includes:
[0035] Perform segmentation processing on the video to be recognized to obtain multiple video segments;
[0036] Extract the image features of each video frame in the multiple video segments;
[0037] Overlay the image features of each video frame in a single video segment to obtain the video sub-feature of the single video segment;
[0038] Arrange the video sub-features corresponding to each video segment in time sequence to obtain the visual features.
[0039] According to one aspect of the present disclosure, there is provided a highlight video recognition device, including:
[0040] An extraction module for extracting the visual features and audio features of the video to be recognized. The video to be recognized is segmented into multiple video segments. The visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence;
[0041] A first encoding module for encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features;
[0042] A splicing module, configured to splice the head and tail of the visual coding feature and the audio coding feature to obtain a spliced feature;
[0043] A second coding module, configured to perform feature coding on the spliced feature based on a self-attention mechanism to obtain a coded spliced feature;
[0044] An identification module, configured to identify highlight video segments among the multiple video segments based on the coded spliced feature.
[0045] In a possible implementation manner, a first coding module is configured to extract a first global context feature of each visual sub-feature in the visual feature; fuse each first global context feature with the corresponding visual sub-feature to obtain multiple first visual sub-features as the visual coding feature; extract a second global context feature of each audio sub-feature in the audio feature; fuse each second global context feature with the corresponding audio sub-feature to obtain multiple first audio sub-features as the audio coding feature.
[0046] In a possible implementation manner, the second coding module is configured to extract a third global context feature of each spliced sub-feature in the spliced feature, where the spliced sub-feature is a first visual sub-feature or a first audio sub-feature; fuse each third global context feature with the corresponding spliced sub-feature to obtain a coded spliced feature.
[0047] In a possible implementation manner, the identification code can be used to split the coded spliced feature at the splicing position to obtain a second visual sub-feature and a second audio sub-feature; fuse the second visual sub-feature and the second audio sub-feature corresponding to the same video segment to obtain multiple fused sub-features; determine whether the segment corresponding to the fused sub-feature is a highlight video segment based on the fused sub-feature.
[0048] In a possible implementation manner, the device further includes:
[0049] A fusion module, configured to fuse the global highlight features extracted from the visual feature and the audio feature with the corresponding visual coding feature and audio coding feature respectively to obtain a visual fusion feature and an audio fusion feature;
[0050] The identification module is configured to identify highlight video segments among the multiple video segments based on the coded spliced feature, the visual fusion feature, and the audio fusion feature.
[0051] In a possible implementation manner, the device further includes:
[0052] The global highlight feature extraction module is used to extract the global highlight features in the visual encoding features and audio encoding features respectively based on the cross-attention mechanism and using the global highlight embedding, where the global highlight embedding is a vector obtained through training for globally abstracting and generalizing highlight features.
[0053] In a possible implementation manner, the recognition module includes:
[0054] The first recognition sub-module is used to obtain a first recognition result based on the encoded concatenated features;
[0055] The second recognition sub-module is used to obtain a second recognition result based on the visual fusion features;
[0056] The third recognition sub-module is used to obtain a third recognition result based on the audio fusion features;
[0057] The weighted fusion module is used to perform weighted fusion on the first recognition result, the second recognition result, and the third recognition result to obtain the recognition result of the highlight segment.
[0058] In a possible implementation manner, the extraction module is used to perform segmentation processing on the video to be recognized to obtain a plurality of video segments; extract the image features of each video frame in the plurality of video segments; superimpose the image features of each video frame in a single video segment to obtain the video sub-features of the single video segment; and arrange the video sub-features corresponding to each video segment in time sequence to obtain visual features.
[0059] According to one aspect of the present disclosure, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0060] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0061] In an embodiment of the present disclosure, by extracting the visual features and audio features of a video to be recognized, the video to be recognized is segmented into multiple video segments. The visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence. Respectively encode the visual features and audio features to obtain visual encoded features and audio encoded features; perform head-to-tail splicing on the visual encoded features and audio encoded features to obtain spliced features; perform feature encoding on the spliced features based on the self-attention mechanism to obtain encoded spliced features; based on the encoded spliced features, identify highlight video segments among the multiple video segments. By encoding the visual features and audio features of the video to be recognized to obtain visual encoded features and audio encoded features, then performing head-to-tail splicing on the visual encoded features (visual sub-features arranged in time sequence) and audio encoded features (audio sub-features arranged in time sequence) to obtain spliced features, and performing feature encoding on the spliced features through the self-attention mechanism to obtain encoded spliced features, it is possible to effectively map the audio and video modality features to the same feature space, and the features of the two different modalities gradually establish a correlation on the same distribution, effectively alleviating the problem of domain misalignment between audio-visual features and creating favorable conditions for the fusion of different modality information. Then, based on the encoded spliced features, identifying the highlight video segments among the multiple video segments can improve the accuracy of the identified highlight video segments.
[0062] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0064] Figure 1 A flowchart showing a method for identifying highlight video segments according to an embodiment of the present disclosure.
[0065] Figure 2 A schematic diagram of an application scenario showing an embodiment of the present disclosure.
[0066] Figure 3 A block diagram showing a device for identifying highlight video segments according to an embodiment of the present disclosure.
[0067] Figure 4 A block diagram showing an electronic device according to an embodiment of the present disclosure.
[0068] Figure 5A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0069] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0070] The special term "exemplary" herein means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.
[0071] The term "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent any one or more elements selected from the set composed of A, B, and C.
[0072] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0073] In recent years, with the large development and application of deep learning in computer vision, many past works have used metric learning methods based on pairs to model the specular highlight detection task. This approach is based on an assumption that there are relatively obvious distinguishable appearances between specular highlight segments and background segments. And based on this assumption, a convolutional neural network is used to encode the video input, and finally the predicted scores of the corresponding segments are output, and the paired scores are strictly required to satisfy the condition: the specular highlight segment score is higher than the background segment score. However, this method does not fully utilize the rich semantic feature connections between segments.
[0074] Studying how to use the information between modalities for complementarity is a good idea. The earliest method to study this idea used simple feature splicing as a modality fusion means. Later work found that using the cross-modal cross-attention method can effectively blend the information of different modalities.
[0075] However, all previous methods have defects to varying degrees. These methods are suboptimal for exploiting the complex relationship between modalities because they are based on an assumption that the feature distributions between multiple modalities are synchronized. But obviously, this assumption does not hold true in practice due to noisy noise interference and the differences in visual and audio features. This also results in the low accuracy of methods developed based on this assumption in video highlight detection tasks.
[0076] In an embodiment of the present disclosure, by extracting visual features and audio features of a video to be identified, the video to be identified is divided into multiple video clips, the visual features include visual sub-features of the multiple video clips arranged in time sequence, and the audio features include audio sub-features of the multiple video clips arranged in time sequence; the visual features and audio features are encoded respectively to obtain visual encoding features and audio encoding features; the visual encoding features and audio encoding features are spliced head to tail to obtain splicing features; feature encoding is performed on the splicing features based on a self-attention mechanism to obtain encoded splicing features; based on the encoded splicing features, highlight video clips among the multiple video clips are identified. By encoding the visual features and audio features of the video to be identified, visual coding features and audio coding features are obtained, and then the visual coding features (time-sequentially arranged visual sub-features) and audio coding features (time-sequentially arranged audio sub-features) are spliced end to end to obtain spliced features. The spliced features are feature encoded through a self-attention mechanism to obtain encoded spliced features. This can effectively map the audio and video modal features to the same feature space, and the features of the two different modalities gradually establish correlation on the same distribution, effectively alleviating the problem of domain misalignment between audio and visual features, and creating favorable conditions for the fusion of different modal information. Then, based on the encoded splicing features, the highlight video segments among the multiple video segments can be identified, which can improve the accuracy of the identified highlight video segments.
[0077] In one possible implementation, the highlight video recognition method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0078] For ease of description, in one or more embodiments of this specification, the execution subject of the highlight video recognition method may be a server. In the following, taking the execution subject as the server as an example, the implementation manners of this method will be introduced. It can be understood that taking the execution subject of this method as the server is only an exemplary illustration and should not be construed as a limitation on this method.
[0079] Figure 1 The flowchart showing the highlight video recognition method according to an embodiment of the present disclosure is as Figure 1 shown, and the highlight video recognition method includes:
[0080] In step S11, extract the visual features and audio features of the video to be recognized;
[0081] The video to be recognized is segmented into multiple video segments. The visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence;
[0082] The video to be recognized can be any video to be recognized whether it is a highlight video segment. During the recognition process, the video to be recognized can be segmented into multiple video segments, and then it is recognized whether the video segments are highlight video segments. The highlight video segments can be wonderful segments or interesting segments that appear in the video.
[0083] The lengths of the multiple video segments segmented here can be the same. For example, the lengths of the video segments are all 50 frames, or the lengths of the video segments are all 100 frames; the lengths of the multiple video segments can also be different, and the present disclosure does not limit this.
[0084] The video to be recognized can be a video file stored in the local storage space of the terminal. Then, the video to be recognized can be read from the local storage space of the terminal and segmented. For example, it can be in the video recording of a local sports event, or it can be the video frames in the local stored shopping mall management video.
[0085] Alternatively, the video to be recognized can also be a video obtained by a real-time acquisition of an image acquisition device. For example, it can be a live video of a sports event, or it can be a video obtained by a real-time acquisition of an image acquisition device located at the entrance of a shopping mall.
[0086] In a possible implementation manner, the extracting the visual features and audio features of the video to be recognized includes: performing segmentation processing on the video to be recognized to obtain multiple video segments; extracting the image features of each video frame in the multiple video segments; superimposing the image features of each video frame in a single video segment to obtain the video sub-feature of the single video segment; arranging the video sub-features corresponding to each video segment in time sequence to obtain the visual features.
[0087] The video to be recognized can be expressed as V = {v t} T t=1 , where v t is a video segment, T is the number of video segments, and t is the time. For consecutive video segments sampled from the video to be recognized, the visual features F v of the video can be extracted by a trained visual feature extraction network (Inflated 3D ConvNet, I3D), and the audio features F of the video can be extracted by a trained audio network (PretrainedAudio Neural a Networks, PANN).
[0088] The visual sub-features and audio sub-features of each video segment are flattened into a feature vector, and then transformed into the same embedding space by the linear layer of the neural network respectively. Therefore, the visual features and audio features of the entire video are respectively expressed as F v = {f1 v ,..., f T v} ∈ R T×d and F a = {f1 a ,..., f T a} ∈ R T×d , where the two dimensions of the embedding space R are distributed as the time series dimension T and the feature dimension d. The specific value of the time series dimension T is the number of video segments obtained by segmentation, and the feature dimension can be 256, for example. f T v is the visual sub-feature, and f T a is the audio sub-feature.
[0089] In this implementation manner, by extracting the image features of each video frame in the multiple video segments and superimposing the image features of each video frame in a single video segment, the video sub-features of a single video segment are obtained. Thus, the obtained video sub-features can more accurately represent all video frames in the video segment. Then, the video sub-features corresponding to each video segment are arranged in time series to obtain visual features, which is convenient for subsequent encoding of the visual features through the attention mechanism based on the context of the video sub-features, so as to accurately extract the key information in the visual features and improve the accuracy of highlight video segment recognition.
[0090] In step S12, the visual features and audio features are encoded respectively to obtain visual encoded features and audio encoded features;
[0091] After obtaining the visual features and audio features, they can be encoded to further extract the key information in the visual features and audio features. For example, extracting the global hypernym and hyponym features in the visual features, and extracting the global hypernym and hyponym features in the audio features. In addition, other features other than the global hypernym and hyponym features can also be extracted, which is not limited in this disclosure.
[0092] In a possible implementation manner, the encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features includes: extracting a first global context feature of each visual sub-feature in the visual features; fusing each of the first global context features with the corresponding visual sub-feature to obtain a plurality of first visual sub-features as the visual encoded features; extracting a second global context feature of each audio sub-feature in the audio features; fusing each of the second global context features with the corresponding audio sub-feature to obtain a plurality of first audio sub-features as the audio encoded features.
[0093] The first global context feature here is used to characterize the context association information between the visual sub-features in the visual features. Fusing the first global context feature with the visual sub-feature can enhance the feature representation of the visual sub-feature. Thus, the obtained visual encoded features can improve the accuracy of the final highlight video segment recognition.
[0094] The second global context feature here is used to characterize the context association information between the audio sub-features in the audio features. Fusing the second global context feature with the audio sub-feature can enhance the feature representation of the audio sub-feature. Thus, the obtained audio encoded features can improve the accuracy of the final highlight video segment recognition.
[0095] In a possible implementation manner, the visual features can be encoded based on the self-attention mechanism to extract the context association information between multiple visual sub-features in the visual features to obtain a plurality of first visual sub-features as the visual encoded features; similarly, the audio features can also be encoded based on the self-attention mechanism to extract the context association information between multiple audio sub-features in the audio features to obtain a plurality of first audio sub-features as the audio encoded features.
[0096] Specifically, for the input of the i-th layer of the encoder First, the input feature is converted into three parts, namely the query Q v , the key K v , and the value S v, The specific conversion process can be seen in formula (1). Then, the first global context feature is captured through the multi-head dot product attention mechanism, and the first global context feature is fused with the visual sub-feature. Specifically, it can be seen in formula (2). Then, a feed-forward network (FFN) is used for transformation and non-linear processing. Specifically, it can be seen in formula (3).
[0097]
[0098]
[0099]
[0100] Among them, is the visual coding feature of the output of the i-th layer of the encoder, and W i q , and are trainable network parameters.
[0101] Similarly, for the audio coding feature it can also be encoded through the above process, which will not be elaborated here.
[0102] In step S13, the visual coding feature and the audio coding feature are concatenated head-to-tail to obtain a concatenated feature;
[0103] Since the visual feature includes visual sub-features of multiple video segments arranged in time sequence, and the audio feature includes audio sub-features of multiple video segments arranged in time sequence. After encoding the visual feature and the audio feature respectively, the obtained visual coding feature can also be multiple first visual sub-features arranged in time sequence, and the audio coding feature can also be multiple first audio sub-features arranged in time sequence.
[0104] When concatenating the visual coding feature and the audio coding feature head-to-tail, the concatenation position can be the last feature in the time sequence of the visual coding feature and the first feature in the time sequence of the audio coding feature, or the first feature in the time sequence of the visual coding feature and the last feature in the time sequence of the audio coding feature.
[0105] Exemplarily, the visual coding feature can be expressed as F n v ={f1 v′ ,...,f T v′}, where f T v′ is the T-th first visual sub-feature, and the audio coding feature can be expressed as F n a ={f1 a′ ,...,fT a′}, where f T a′ is the T-th first audio sub-feature.
[0106] Then, concatenate the visual encoding feature and the audio encoding feature at the head and tail, and the obtained concatenated feature can be expressed as F av ={f1 v′ ,..., f T v′ , f1 a′ ,..., f T a′}}.
[0107] Obviously, after concatenating the visual encoding feature and the audio encoding feature with the time series dimension length of T at the head and tail, the length of the obtained concatenated feature in the time series dimension is 2T.
[0108] In step S14, perform feature encoding on the concatenated feature based on the self-attention mechanism to obtain the encoded concatenated feature;
[0109] When performing feature encoding on the concatenated feature based on the self-attention mechanism, the concatenated feature will be regarded as a whole, and the sub-features in this whole will be encoded. Through this process, not only the context features between the first visual sub-features and the context features between the first audio sub-features will be extracted, but also the concatenated feature will be regarded as a whole, and the global context features between the first visual sub-features and the first audio sub-features will be extracted from the overall global perspective.
[0110] That is, based on the self-attention mechanism, perform encoding processing on the overall sequence composed of the first visual sub-features and the first audio sub-features to obtain the encoded concatenated feature. For the convenience of description, in the encoded concatenated feature, the feature updated from the first visual sub-feature can be called the second visual sub-feature, and the feature updated from the second visual sub-feature can be called the second audio sub-feature.
[0111] In this way, in the encoded concatenated feature, the second visual sub-feature will contain the context features of the second audio sub-feature; the second audio sub-feature will also contain the context features of the second visual sub-feature.
[0112] In a possible implementation manner, the performing feature encoding on the concatenated feature based on the self-attention mechanism to obtain the encoded concatenated feature includes: extracting the third global context feature of each concatenated sub-feature in the concatenated feature, where the concatenated sub-feature is the first visual sub-feature or the first audio sub-feature; fusing each of the third global context features with the corresponding concatenated sub-feature to obtain the encoded concatenated feature.
[0113] For ease of description, the sub - features in the splicing feature are referred to as splicing sub - features here. Since the splicing feature is obtained by splicing the visual coding feature and the audio coding feature, and the visual coding feature contains the first visual sub - feature, and the audio coding feature contains the first audio sub - feature, the splicing sub - feature is the first visual sub - feature or the first audio sub - feature.
[0114] The third global context feature here is used to characterize the context association information between the splicing sub - features in the splicing feature. Fusing the third global context feature with the corresponding splicing sub - feature can enhance the feature representation of each splicing sub - feature. Thus, the accuracy of the final highlight video segment recognition can be improved.
[0115] Since audio features and visual features are features in two different domains, the domain gap between the two domains is very large, and the manifestation forms of the features in the two domains are different. Maybe at a certain moment in the video, the visual feature is highlight content, but the music may be calm, that is, the audio feature is not highlight content because the audio may be background music, and this background music may have been processed by someone and may be irrelevant to the video's picture effect. Another example is that the video's picture may shake for some reason, and the video may not belong to highlight content, but it may be heard from the audio that something interesting happened at that moment, that is, there is highlight content at that moment. That is to say, the visual feature and the audio feature may not be aligned in terms of the highlight time.
[0116] Then, directly superimposing the visual feature and the audio feature at the same moment to estimate the highlight video segment may not have a good effect.
[0117] In this implementation, after splicing the visual coding feature and the audio coding feature end - to - end, by calculating the global context feature between the splicing sub - features in the overall splicing, some relevant connections between the visual coding feature and the audio coding feature can be captured, enabling the visual coding feature to capture the situation of the audio coding feature and the audio coding feature to capture the situation of the visual coding feature. The domain gap between vision and audio is reduced, and the accuracy of highlight segment recognition is improved.
[0118] In step S15, based on the encoded splicing feature, identify the highlight video segments in the multiple video segments.
[0119] After obtaining the encoded splicing feature, the highlight video segments in the multiple video segments can be identified based on the encoded splicing feature. For example, a trained network can be used to identify the highlight video segments. The specific identification process can refer to the possible implementation provided in this disclosure and will not be elaborated here.
[0120] In an embodiment of the present disclosure, by extracting the visual features and audio features of the video to be recognized, the video to be recognized is segmented into multiple video segments. The visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence. The visual features and audio features are respectively encoded to obtain visual encoded features and audio encoded features. The visual encoded features and audio encoded features are concatenated at the head and tail to obtain concatenated features. The concatenated features are encoded based on the self-attention mechanism to obtain encoded concatenated features. Based on the encoded concatenated features, the highlight video segments in the multiple video segments are recognized. By encoding the visual features and audio features of the video to be recognized to obtain visual encoded features and audio encoded features, then concatenating the visual encoded features (visual sub-features arranged in time sequence) and audio encoded features (audio sub-features arranged in time sequence) at the head and tail to obtain concatenated features, and encoding the concatenated features through the self-attention mechanism to obtain encoded concatenated features, it is possible to effectively map the audio and video modality features to the same feature space, and the features of the two different modalities gradually establish a correlation on the same distribution, effectively alleviating the problem of domain misalignment between the audio and visual features, creating favorable conditions for the fusion of different modality information. Then, based on the encoded concatenated features, recognizing the highlight video segments in the multiple video segments can improve the accuracy of the recognized highlight video segments.
[0121] In a possible implementation manner, the recognizing the highlight video segments in the multiple video segments based on the encoded concatenated features includes: at the concatenation position, splitting the encoded concatenated features to obtain second visual sub-features and second audio sub-features; fusing the second visual sub-features and second audio sub-features corresponding to the same video segment to obtain multiple fused sub-features; and based on the fused sub-features, determining whether the segment corresponding to the fused sub-feature is a highlight video segment.
[0122] In the encoded concatenated features, the second visual sub-features and second audio sub-features capture some correlation information of each other. When recognizing the highlight video segments, the encoded concatenated features can be split. The splitting position is the concatenation position in step S13. The two features obtained after splitting are respectively composed of the second visual sub-features and second audio sub-features. These two features have the same length in the time dimension, both being T.
[0123] For the two features after disassembly, the second visual sub-feature and the second audio sub-feature corresponding to the same video segment can be fused to obtain a fused sub-feature. A single fused sub-feature contains both the visual feature (the second visual sub-feature) of the video segment and the audio feature (the second audio sub-feature) of the video segment. The fusion here can specifically be adding the second visual sub-feature and the second audio sub-feature of the same video segment.
[0124] When determining whether the segment corresponding to the fused sub-feature is a highlight video segment based on the fused sub-feature, the fused sub-feature can be input into the trained network classification layer to obtain the result of whether the segment is a highlight video segment. Its essence can be a binary classification network for determining whether the input feature is a highlight video segment, or alternatively, it can output a score, which is a value between 0 and 1. The higher the score, the greater the likelihood that the video segment is a highlight video segment. For ease of description, the recognition result output based on the encoded splicing feature is referred to as the first recognition result here.
[0125] In a possible implementation manner, after encoding the visual feature and the audio feature respectively to obtain the visual encoded feature and the audio encoded feature, the method further includes: fusing the global highlight features extracted from the visual feature and the audio feature with the corresponding visual encoded feature and audio encoded feature respectively to obtain a visual fused feature and an audio fused feature; the recognizing the highlight video segments in the multiple video segments based on the encoded splicing feature includes: recognizing the highlight video segments in the multiple video segments based on the encoded splicing feature, the visual fused feature, and the audio fused feature.
[0126] In a possible implementation manner, the method for extracting the global highlight features from the visual feature and the audio feature includes: based on the cross-attention mechanism, using the global highlight embedding to extract the global highlight features in the visual encoded feature and the audio encoded feature respectively, where the global highlight embedding is a vector obtained through training for globally abstracting and generalizing the highlight features.
[0127] The global highlight embedding is a vector obtained through training for globally abstracting and generalizing the highlight features. For the specific training process, reference can be made to the possible implementation manners provided in this disclosure, which will not be elaborated here. The global highlight embedding can be considered as the common features of the highlight features in various videos after abstracting and generalizing the highlight features in various videos.
[0128] Taking the visual encoded feature as an example, in the process of using the global highlight embedding to extract the global highlight feature of the visual encoded feature, the visual encoded feature and the global highlight embedding G as the input of the cross-attention mechanism. Specifically, G can be used as the query, and the visual encoding features are regarded as the values, querying and aggregating the visual context information in the visual encoding features to obtain the global highlight feature G v . Then, the visual encoding features and the global highlight feature G v are summed to obtain the visual fusion feature This process can be expressed as
[0129] Taking the audio encoding features as an example, in the process of using the global highlight embedding to extract the global highlight features of the audio encoding features, the audio encoding features and the global highlight embedding G can be used as the input of the cross-attention mechanism. Specifically, G can be used as the query, and the audio encoding features are regarded as the values, querying and aggregating the visual context information in the audio encoding features to obtain the global highlight feature G a . Then, the audio encoding features and the global highlight feature G a are summed to obtain the audio fusion feature This process can be expressed as
[0130] In a possible implementation, the global highlight embedding G is obtained through training. G is initially initialized with a value. During training, the initialized G and after being operated by the cross-attention mechanism, G' is obtained. Then, G is updated based on the gradient of G', and the next cross-attention operation is continued, continuously iterating to update G.
[0131] In this implementation, by using the global highlight embedding based on the cross-attention mechanism, the global highlight features in the visual encoding features and audio encoding features are respectively extracted, and the global highlight features are respectively fused with the corresponding visual encoding features and audio encoding features to obtain the visual fusion feature and audio fusion feature. Since the global highlight embedding is a vector obtained through training for globally abstracting and generalizing the highlight features, therefore, by using the cross-attention mechanism to abstract and generalize the single-modal sequence into a global feature, this global feature can be considered as the highlight feature in the visual encoding features and audio encoding features. Thus, the obtained visual fusion feature and audio fusion feature can greatly enhance the feature representation ability and improve the accuracy of the recognized highlight video segments.
[0132] Since the visual fusion features and the audio fusion features respectively represent the features of the visual and audio modalities, and the context information between the audio-visual modalities is fully exploited in the encoded concatenated features, therefore, identifying the highlight video segments in multiple video segments based on the encoded concatenated features, the visual fusion features, and the audio fusion features can improve the accuracy of the identified highlight video segments.
[0133] In a possible implementation manner, the identifying the highlight video segments in the multiple video segments based on the encoded concatenated features, the visual fusion features, and the audio fusion features includes: obtaining a first identification result based on the encoded concatenated features; obtaining a second identification result based on the visual fusion features; obtaining a third identification result based on the audio fusion features; and performing weighted fusion on the first identification result, the second identification result, and the third identification result to obtain an identification result of the highlight segment.
[0134] Obtaining a first identification result based on the encoded concatenated features The process can refer to the relevant descriptions above and will not be elaborated here.
[0135] Correspondingly, the visual fusion features and the audio fusion features are also respectively input into the trained network classification layer to obtain a second identification result and a third identification result In an example, and can be a score, which is a value between 0 and 1. The higher the score, the greater the possibility that the video segment is a highlight video segment. Then, the three scores and can be weighted and summed to obtain a weighted sum score, and then based on the weighted sum score, it is determined whether the video segment is a highlight video segment. The weights of the three scores can be the same or different, and the present disclosure does not limit this.
[0136] In another example, and can also be the results of whether the video segment is a highlight video segment. In the case of giving priority to accuracy, when and both indicate that the video segment is a highlight segment, it can be determined that the video segment is a highlight video segment.
[0137] In addition, it is also possible to first perform weighted fusion on the encoded concatenated features, the visual fusion features, and the audio fusion features, and then identify the weighted fusion features to obtain the highlight video segments in multiple video segments. The present disclosure does not elaborate on this.
[0138] The implementation of a training process of the present disclosure will be described below. The highlight video recognition method provided by the present disclosure can be implemented based on a recognition network. The input of the recognition network can be multiple segmented video clips, and the output is the recognition result of the highlight video. Then, during training, each video clip can be labeled, and the expected output y of each video clip can be labeled. Then, based on the difference between the expected output y and the actual output result, the parameters in the recognition network can be updated by gradient. During this process, the global highlight embedding G will be updated by gradient at the same time. For the update process of G, reference can be made to the relevant description above.
[0139] Please refer to Figure 2 , which is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. As Figure 2 shown, through the visual feature extraction network E v the visual features of the video clip are extracted to obtain visual features F v , and through the audio feature extraction network E a the audio features of the video clip are extracted to obtain audio features F a . Figure 2 Among them, each small grid of the visual feature F v represents the visual sub-feature of a video clip, and each small grid in the audio feature F a represents the audio sub-feature of a video clip.
[0140] Then, through the intra-modal encoding module, the self-attention mechanism is used to encode the visual feature F v and the audio feature F a respectively to obtain the visual encoded feature and the audio encoded feature Then, based on the cross-attention mechanism, the global highlight embedding G is used to extract the global highlight features in the visual encoded feature and the audio encoded feature respectively, and the global highlight features are fused with the corresponding visual encoded feature and audio encoded feature to obtain the visual fusion feature and the audio fusion feature
[0141] Then, through the co-occurrence encoding module, and are concatenated at the head and tail, and then encoded based on self-attention to obtain Then, at the splicing position, is split and then spliced to obtain
[0142] Respectively for and Perform highlight video segment estimation to identify highlight video segments in multiple video segments. For the specific estimation process, refer to the relevant descriptions above. Details are not elaborated here.
[0143] It can be understood that for the above-mentioned various method embodiments mentioned in the present disclosure, without violating the principle logic, they can be combined with each other to form combined embodiments. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0144] In addition, the present disclosure also provides a highlight video recognition device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any one of the highlight video recognition methods provided by the present disclosure. For the corresponding technical solutions and descriptions, refer to the corresponding records in the method part and will not be elaborated further.
[0145] Figure 3 The block diagram of the highlight video recognition device according to an embodiment of the present disclosure is shown, as Figure 3 shown, the device 30 includes:
[0146] An extraction module 31, configured to extract visual features and audio features of a video to be recognized. The video to be recognized is segmented into multiple video segments. The visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence;
[0147] A first encoding module 32, configured to encode the visual features and audio features respectively to obtain visual encoded features and audio encoded features;
[0148] A splicing module 33, configured to perform head-to-tail splicing on the visual encoded features and audio encoded features to obtain spliced features;
[0149] A second encoding module 34, configured to perform feature encoding on the spliced features based on a self-attention mechanism to obtain encoded spliced features;
[0150] A recognition module 35, configured to identify highlight video segments in the multiple video segments based on the encoded spliced features.
[0151] In a possible implementation manner, the first encoding module is configured to extract first global context features of each visual sub-feature in the visual features; fuse each first global context feature with the corresponding visual sub-feature to obtain multiple first visual sub-features as the visual encoded features; extract second global context features of each audio sub-feature in the audio features; fuse each second global context feature with the corresponding audio sub-feature to obtain multiple first audio sub-features as the audio encoded features.
[0152] In a possible implementation, the second encoding module is configured to extract third global context features of each splicing sub-feature in the splicing feature, where the splicing sub-feature is a first visual sub-feature or a first audio sub-feature; and fuse each of the third global context features with the corresponding splicing sub-feature to obtain an encoded splicing feature.
[0153] In a possible implementation, the identification code is configured to split the encoded splicing feature at the splicing position to obtain a second visual sub-feature and a second audio sub-feature; fuse the second visual sub-feature and the second audio sub-feature corresponding to the same video segment to obtain a plurality of fused sub-features; and determine whether the segment corresponding to the fused sub-feature is a highlight video segment based on the fused sub-features.
[0154] In a possible implementation, the apparatus further includes:
[0155] A fusion module, configured to fuse the global highlight features extracted from the visual feature and the audio feature with the corresponding visual encoding feature and audio encoding feature respectively to obtain a visual fusion feature and an audio fusion feature;
[0156] The identification module is configured to identify highlight video segments in the plurality of video segments based on the encoded splicing feature, the visual fusion feature, and the audio fusion feature.
[0157] In a possible implementation, the apparatus further includes:
[0158] A global highlight feature extraction module, configured to extract global highlight features in the visual encoding feature and the audio encoding feature respectively based on a cross-attention mechanism and using a global highlight embedding, where the global highlight embedding is a vector obtained through training for globally abstracting and generalizing highlight features.
[0159] In a possible implementation, the identification module includes:
[0160] A first identification sub-module, configured to obtain a first identification result based on the encoded splicing feature;
[0161] A second identification sub-module, configured to obtain a second identification result based on the visual fusion feature;
[0162] A third identification sub-module, configured to obtain a third identification result based on the audio fusion feature;
[0163] A weighted fusion module, configured to perform weighted fusion on the first identification result, the second identification result, and the third identification result to obtain an identification result of the highlight segment.
[0164] In a possible implementation, the extraction module is configured to segment the video to be recognized to obtain a plurality of video segments; extract the image features of each video frame in the plurality of video segments; superimpose the image features of each video frame in a single video segment to obtain a video sub-feature of the single video segment; and arrange the video sub-features corresponding to the respective video segments in time sequence to obtain visual features.
[0165] This method has a specific technical association with the internal structure of a computer system and can solve technical problems of how to improve the hardware operation efficiency or execution effect (including reducing the data storage amount, reducing the data transmission amount, increasing the hardware processing speed, etc.), thereby obtaining a technical effect of improving the internal performance of the computer system in line with natural laws.
[0166] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0167] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0168] The embodiments of the present disclosure also propose an electronic device, including: a processor; and a memory for storing instructions executable by the processor. Wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.
[0169] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above methods.
[0170] The electronic device can be provided as a terminal, a server, or other forms of devices.
[0171] Figure 4 A block diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. For example, the electronic device 800 can be a terminal device such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.
[0172] Refer to Figure 4, the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0173] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0174] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0175] The power component 806 provides power to the various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0176] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0177] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0178] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0179] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or a charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0180] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 may access a communication standard-based wireless network, such as a wireless network (Wi-Fi), a second-generation mobile communication technology (2G), a third-generation mobile communication technology (3G), a fourth-generation mobile communication technology (4G), a long-term evolution of the universal mobile communication technology (LTE), a fifth-generation mobile communication technology (5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0181] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0182] In an exemplary embodiment, a non-volatile computer-readable storage medium is further provided, such as a memory 804 including computer program instructions, and the computer program instructions may be executed by a processor 820 of the electronic device 800 to complete the above method.
[0183] The present disclosure relates to the field of augmented reality. By obtaining the image information of a target object in the real environment, relevant features, states, and attributes of the target object are detected or recognized through various vision-related algorithms, so as to obtain an AR effect combining virtual and real that matches a specific application. Exemplarily, the target object may involve the face, limbs, gestures, actions, etc. related to the human body, or identification marks, markers related to objects, or sand tables, display areas, or display items related to venues or places. The vision-related algorithms may involve visual positioning, SLAM, three-dimensional reconstruction, image registration, background segmentation, key point extraction and tracking of objects, pose or depth detection of objects, etc. The specific application can not only involve interactive scenarios such as guiding, navigation, explanation, reconstruction, virtual effect overlay display related to real scenes or items, but also involve special effect processing related to people, such as makeup beautification, limb beautification, special effect display, virtual model display, etc. The relevant features, states, and attributes of the target object can be detected or recognized through a convolutional neural network. The convolutional neural network is a network model obtained by training the model based on a deep learning framework.
[0184] Figure 5 FIG. 4 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 5 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0185] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphics user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.
[0186] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.
[0187] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0188] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0189] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0190] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0191] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0192] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0193] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0195] The computer program product may be implemented specifically by means of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), and so on.
[0196] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they are not elaborated herein again.
[0197] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0198] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0199] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A high-brightness video recognition method, characterized in that, Including: Extracting visual features and audio features of the video to be recognized, where the video to be recognized is segmented into multiple video segments, the visual features include visual sub-features of multiple video segments arranged in time sequence, and the audio features include audio sub-features of multiple video segments arranged in time sequence; Encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features; Performing head-to-tail splicing on the visual encoded features and audio encoded features to obtain spliced features; Performing feature encoding on the spliced features based on the self-attention mechanism to obtain encoded spliced features; Identifying highlight video segments among the multiple video segments based on the encoded spliced features.
2. The method according to claim 1, wherein Encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features, including: Extracting the first global context feature of each visual sub-feature in the visual features; Fusing each of the first global context features with the corresponding visual sub-feature to obtain multiple first visual sub-features as the visual encoded features; Extracting the second global context feature of each audio sub-feature in the audio features; Fusing each of the second global context features with the corresponding audio sub-feature to obtain multiple first audio sub-features as the audio encoded features.
3. The method according to any one of claims 1 or 2, characterized in that The performing feature encoding on the spliced features based on the self-attention mechanism to obtain encoded spliced features includes: Extracting the third global context feature of each spliced sub-feature in the spliced features, where the spliced sub-feature is a first visual sub-feature or a first audio sub-feature; Fusing each of the third global context features with the corresponding spliced sub-feature respectively to obtain encoded spliced features.
4. The method according to claim 3, characterized in that The identifying highlight video segments among the multiple video segments based on the encoded spliced features includes: At the splicing position, splitting the encoded spliced features to obtain second visual sub-features and second audio sub-features; Fusing the second visual sub-features and second audio sub-features corresponding to the same video segment to obtain multiple fused sub-features; Based on the fused sub-features, determining whether the segment corresponding to the fused sub-feature is a highlight video segment.
5. The method according to any one of claims 1-4, characterized in that, After encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features, the method further includes: Fusing the global highlight features extracted from the visual features and audio features with the corresponding visual encoded features and audio encoded features respectively to obtain visual fused features and audio fused features; The identifying highlight video segments among the multiple video segments based on the encoded spliced features includes: Identifying highlight video segments among the multiple video segments based on the encoded spliced features, the visual fused features and audio fused features.
6. The method according to claim 5, wherein The method for extracting global highlight features from the visual features and audio features includes: Based on the cross-attention mechanism, using global highlight embeddings to extract global highlight features in the visual encoded features and audio encoded features respectively, where the global highlight embeddings are vectors obtained through training for globally abstracting and generalizing highlight features.
7. The method according to claim 5, characterized in that, Identifying the highlight video segments in the multiple video segments based on the encoded splicing features, the visual fusion features, and the audio fusion features includes: Obtaining a first recognition result based on the encoded splicing features; Obtaining a second recognition result based on the visual fusion features; Obtaining a third recognition result based on the audio fusion features; Performing weighted fusion on the first recognition result, the second recognition result, and the third recognition result to obtain the recognition result of the highlight segments.
8. The method according to claim 1, characterized in that Extracting the visual features and audio features of the video to be recognized includes: Performing segmentation processing on the video to be recognized to obtain multiple video segments; Extracting the image features of each video frame in the multiple video segments; Overlaying the image features of each video frame in a single video segment to obtain the video sub-features of the single video segment; Arranging the video sub-features corresponding to each video segment in time sequence to obtain visual features.
9. A high-brightness video recognition device, characterized in that, Including: An extraction module for extracting the visual features and audio features of the video to be recognized, the video to be recognized being segmented into multiple video segments, the visual features including visual sub-features of multiple video segments arranged in time sequence, and the audio features including audio sub-features of multiple video segments arranged in time sequence; A first encoding module for encoding the visual features and audio features respectively to obtain visual encoded features and audio encoded features; A splicing module for splicing the visual encoded features and audio encoded features head to tail to obtain splicing features; A second encoding module for performing feature encoding on the splicing features based on the self-attention mechanism to obtain the encoded splicing features; A recognition module for identifying the highlight video segments in the multiple video segments based on the encoded splicing features.
10. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1 to 8.