A Method and Device for Multi-Scale Two-Stream Attention Video-Language Event Prediction
Through the multi-scale dual-stream attention video language event prediction method, multi-scale video features are generated and fused with subtitles and candidate event features, the problem of low accuracy of existing video prediction models is solved and higher event prediction accuracy is achieved.
Patent Information
- Application Number
- CN202210412836.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing video prediction model has low accuracy and cannot meet the actual usage requirements. In particular, deep semantic understanding technologies in video Q&A and video prediction have not yet been widely used.
The multi-scale dual-stream attention video language event prediction method is adopted, and multi-scale video features are generated through the multi-scale video processing module. Combined with the dual-stream cross-modal fusion module and the event prediction module, the appearance and action features of the video frame are processed respectively, and fused with the subtitle features and future candidate event features to generate fusion features of different scales, and finally determine the event prediction results.
It improves the accuracy of video prediction, reduces redundant features, avoids mutual interference between different modes, and enhances the accuracy of event prediction.
Smart Images

Figure CN115019137B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method and device for multi-scale two-stream attention video-language event prediction. Background Art
[0002] In recent years, the rapid development of the Internet has triggered an information explosion, making the current era also known as the information age. As the most important and densest carrier of information, videos have become very common on the network. Analyzing such a vast amount of data that is closely related to people's daily lives can generate huge value and even bring about major social changes. Some video analysis technologies have been applied in social life, such as intelligent review of bad video content, video object detection, and video face recognition. However, related research technologies on deep video semantic understanding, represented by video question answering and video prediction, have not been widely applied on a large scale. One of the reasons is that the performance of existing models is too poor to meet the actual usage requirements. Among them, video prediction is to predict future candidate events based on video semantic understanding.
[0003] Therefore, how to improve the accuracy of video prediction is an urgent problem to be solved at present. Summary of the Invention
[0004] The present invention provides a method and device for multi-scale two-stream attention video-language event prediction, which are used to solve the defect of low accuracy of video prediction in the prior art and achieve the improvement of the accuracy of video prediction.
[0005] The present invention provides a method for multi-scale two-stream attention video-language event prediction, including: obtaining original input data; wherein, the original input data includes a target video stream, subtitles corresponding to the target video stream, and a plurality of future candidate events; inputting the original input data into a multi-scale two-stream attention video-language event prediction model to obtain an event prediction result of the target video stream; wherein, the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module; the multi-scale video processing module is used to generate multi-scale video features based on video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitles, the features of the plurality of future candidate events, and the multi-scale video features; the event prediction module is used to respectively obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales, and determine a final event prediction result of the target video stream based on the event prediction results.
[0006] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, the generation of the multi-scale video features includes:
[0007] Sampling the target video stream with different sampling steps to obtain video frames with different sampling scales;
[0008] Performing feature extraction on the video frames with different sampling scales to obtain multi-scale video features.
[0009] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, the video frames with different sampling scales include: video frames with a dense sampling scale, video frames with a general sampling scale, and video frames with a sparse sampling scale; correspondingly, the performing feature extraction on the video frames with different sampling scales to obtain multi-scale video features includes:
[0010] Based on the video frames with the dense sampling scale and the pre-trained SlowFast model, obtaining the first video feature of the video frames with the dense sampling scale;
[0011] Based on the video frames with the general sampling scale and the pre-trained ResNet-152 model, obtaining the second video feature of the video frames with the general sampling scale;
[0012] Based on the video frames with the sparse sampling scale and the pre-trained SlowFast model, obtaining the third video feature of the video frames with the sparse sampling scale; based on the video frames with the sparse sampling scale and the pre-trained ResNet-152 model, obtaining the fourth video feature of the video frames with the sparse sampling scale; and splicing the third video feature and the fourth video feature to obtain the fifth video feature;
[0013] Determining multi-scale video features based on the first video feature, the second video feature, and the fifth video feature.
[0014] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, the generation of the first fusion video features with different scales includes the following steps:
[0015] Based on the single-modal feature conversion layer guided by future candidate events, respectively fusing the video features with different scales in the multi-scale video features with the features of each future candidate event to obtain the sixth video feature of the video features with different scales guided by future candidate events;
[0016] Based on the cross-modal fusion layer of dual-stream video subtitles, fuse the video features of different scales in the multi-scale video features with the features of the subtitles corresponding to the target video stream respectively, and concatenate the fused features with the features of each of the future candidate events to obtain video features of different scales guided by subtitles; and input the video features of different scales guided by subtitles into the single-modal feature conversion layer guided by the future candidate events to obtain the seventh video feature of the video features of each scale;
[0017] Concatenate the sixth video feature and the seventh video feature corresponding to the video features of each scale to obtain the first fused video feature of each scale.
[0018] According to a method for multi-scale dual-stream attention video-language event prediction provided by the present invention, the generation of the first fused subtitle features of different scales includes the following steps:
[0019] Based on the single-modal feature conversion layer guided by future candidate events, fuse the features of the subtitles corresponding to the target video stream with the features of each of the future candidate events respectively to obtain the first subtitle features guided by future candidate events;
[0020] Based on the cross-modal fusion layer of dual-stream video subtitles, fuse the features of the subtitles corresponding to the target video stream with the multi-scale video features respectively to obtain subtitle features guided by video frames of different scales; and based on the single-modal feature conversion layer guided by the future candidate events, fuse the fused features with the features of each of the future candidate events respectively to obtain multiple second subtitle features guided by the video;
[0021] Concatenate the multiple first subtitle features and the multiple second subtitle features to obtain the first fused subtitle features.
[0022] According to a method for multi-scale dual-stream attention video-language event prediction provided by the present invention, the multi-scale dual-stream attention video-language event prediction model further includes a subtitle and future candidate event feature extraction module. Correspondingly, the features of the subtitle and the features of the multiple future candidate events are generated based on the subtitle and future candidate event feature extraction module, including:
[0023] Input the subtitle corresponding to the target video stream into the subtitle and future candidate event feature extraction module to obtain the features of the subtitle;
[0024] Input the multiple future candidate events into the subtitle and future candidate event feature extraction module to obtain the features of the multiple future candidate events.
[0025] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, the multi-scale two-stream attention video-language event prediction model further includes a multi-scale fusion module, and the multi-scale fusion module is used to fuse the first fusion video features of different scales to obtain second fusion video features, and is used to fuse the first fusion caption features of different scales to obtain second fusion caption features.
[0026] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, obtaining the future candidate event prediction result of the target video stream based on the first fusion video feature and the first fusion caption feature, includes:
[0027] Compress the second fusion video feature to obtain the compressed second fusion video feature; and compress the second fusion caption feature to obtain the compressed second fusion caption feature;
[0028] Perform event prediction based on the compressed second fusion video feature to obtain multiple first scores corresponding to multiple future candidate events of the target video stream; and perform event prediction based on the compressed second fusion caption feature to obtain multiple second scores corresponding to multiple future candidate events of the target video stream;
[0029] Add the first score of each future candidate event to the second score of each future candidate event to obtain the total score corresponding to each future candidate event of the target video stream;
[0030] Determine the future candidate events corresponding to the target video stream based on the total score corresponding to each future candidate event of the target video stream.
[0031] A method for multi-scale two-stream attention video-language event prediction provided by the present invention, determining the multi-scale video feature based on the first video feature, the second video feature, and the fifth video feature, includes:
[0032] Convert the first video feature, the second video feature, and the fifth video feature into the same dimension;
[0033] Based on the Transformer encoder, perform temporal encoding on the first video feature, the second video feature, and the fifth video feature after dimension conversion respectively to obtain the encoded first video feature, second video feature, and fifth video feature;
[0034] Use the encoded first video feature, second video feature, and fifth video feature as the multi-scale video feature.
[0035] The present invention also provides a device for multi-scale two-stream attention video-language event prediction, including:
[0036] An acquisition module for acquiring original input data; wherein, the original input data includes a target video stream, subtitles corresponding to the target video stream, and a plurality of future candidate events;
[0037] A processing module for inputting the original input data into a multi-scale two-stream attention video-language event prediction model to obtain an event prediction result of the target video stream;
[0038] Wherein, the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module;
[0039] The multi-scale video processing module is used to generate multi-scale video features based on video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features at different scales and first fusion subtitle features at different scales based on the features of the subtitles, the features of the plurality of future candidate events, and the multi-scale video features; the event prediction module is used to obtain event prediction results based on the first fusion video features at different scales and the first fusion subtitle features at different scales respectively, and determine the final event prediction result of the target video stream based on the event prediction results.
[0040] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the method for multi-scale two-stream attention video-language event prediction as described in any one of the above are implemented.
[0041] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for multi-scale two-stream attention video-language event prediction as described in any one of the above are implemented.
[0042] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method for multi-scale two-stream attention video-language event prediction as described in any one of the above are implemented.
[0043] The method and device for multi-scale two-stream attention video-language event prediction provided by the present invention obtain multi-scale video features by performing multi-scale processing on a video, making the extracted video features more reasonable, and based on the multi-scale video features, fusing with the features of subtitles and the features of multiple future candidate events to obtain first fusion video features of different scales and first fusion subtitle features of different scales, and respectively performing event prediction based on the first fusion video features of different scales and the first fusion subtitle features of different scales, and then combining the prediction results to determine the final event prediction result, comprehensively extracting features, reducing redundant features, avoiding the adverse effects caused by mutual interference between different modalities, and effectively improving the accuracy of event prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 is one of the flow diagrams of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0046] Figure 2 is the second flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0047] Figure 3 is the third flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0048] Figure 4 is the fourth flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0049] Figure 5 is the fifth flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0050] Figure 6 is the sixth flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0051] Figure 7 is the seventh flow diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0052] Figure 8 is the framework diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0053] Figure 9 It is a schematic structural diagram of the device for multi-scale two-stream attention video-language event prediction provided by the present invention;
[0054] Figure 10 It is a schematic structural diagram of the electronic device provided by the present invention.
[0055] Reference numerals:
[0056] 1010: Processor; 1020: Communication interface; 1030: Memory; 1040: Communication bus. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0058] For the convenience of understanding, the background technology related to the present invention will be briefly introduced first.
[0059] Video understanding generally requires characterizing the video from two aspects: appearance features and motion features, that is, identifying the actions in the event, their sequence, and the objects in the video shot. Then, cross-modal fusion of video and language is performed. In the current technology, usually the same method is used to process these two different video features. A common approach is to take frames as units and splice the appearance features and motion features of each frame. Subsequent processing makes no difference to the two features. However, compared with the motion features, the appearance features of each frame are relatively easy to extract. Therefore, if the same method is used to extract the two features, it will cause redundancy of the extracted appearance features, which is not conducive to the training and use of the model, and will result in a low accuracy of event prediction.
[0060] In addition, in the current technology, usually a single-stream cross-modal fusion method is often adopted, that is, first obtain the joint representation of the two modalities of video and subtitle, and then generate the prediction score through the prediction module with the joint representation result. However, since some specific samples may only require the information of one modality to answer, blindly using the joint representation vector to predict future candidate events will inevitably introduce redundant information and the situation of mutual interference between different modalities.
[0061] The following will be combined with Figures 1 - 10 Describe the method and device for multi-scale two-stream attention video-language event prediction of the present invention.
[0062] Figure 1 This is one of the schematic flowcharts of the method for multi-scale two-stream attention video-language event prediction provided by the present invention. It can be understood that Figure 1 the method in
[0063] As Figure 1 shown, the method for predicting multi-scale two-stream attention video-language events provided by the present invention includes the following steps:
[0064] Step 110: Obtain the original input data.
[0065] Among them, the original input data includes a target video stream, the subtitle corresponding to the target video stream, and multiple future candidate events.
[0066] Among them, the subtitle corresponding to the target video stream can be: the dialogue text of the target person in the video. The future candidate events can be the events that may occur in the future defined according to the actions being performed by the target person.
[0067] Step 120: Input the original input data into the multi-scale two-stream attention video-language event prediction model to obtain the event prediction result of the target video stream.
[0068] Among them, the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module. The multi-scale video processing module is used to generate multi-scale video features based on the video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitle, the features of the multiple future candidate events, and the multi-scale video features; the event prediction module is used to obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales respectively, and determine the final event prediction result of the target video stream based on the event prediction results.
[0069] It can be understood that when predicting events for a video stream, it is necessary to determine the appearance features and action features in the video frame images of the video stream to predict the upcoming events. Among them, the appearance features can be, for example, the features corresponding to the characters and the scene, and the action features can be, for example, the action features of the characters. And the changes in the characters and the scene in each video frame of a video stream are relatively small compared to the changes in the actions of the characters. Therefore, if the same method is used to extract the corresponding features of the characters, the scene, and the actions of the characters, it will cause redundancy in the extracted appearance features, which is not conducive to the training and use of the model, and will result in a low accuracy of event prediction. Therefore, the present invention adopts a multi-scale video processing module for generating multi-scale video features.
[0070] It can also be understood that since the subtitle features and video features are features of two different modalities, therefore, event prediction is performed on the video stream based on the features of the two different modalities respectively, and the two prediction results are used as the final prediction result. This can not only perform event prediction in a multi-modal manner, but also comprehensively extract features, reduce redundant features, and avoid the adverse effects caused by mutual interference between different modalities.
[0071] In addition, not only based on the subtitle features and the multi-scale video features, but also based on the features of multiple future candidate events, first fusion video features of different scales and first fusion subtitle features of different scales are generated, so that the extracted video features and subtitle features include features related to each future candidate event. When performing event prediction subsequently, the prediction result of the video can be obtained based on the association between the video, subtitle, and candidate event.
[0072] The method for multi-scale two-stream attention video-language event prediction provided by the present invention obtains multi-scale video features by performing multi-scale processing on the video, making the extracted video features more reasonable. Based on the multi-scale video features, the features of the subtitle, and the features of multiple future candidate events, first fusion video features of different scales and first fusion subtitle features of different scales are fused. Event prediction is performed respectively based on the first fusion video features of different scales and the first fusion subtitle features of different scales, and then the prediction results are combined to determine the final event prediction result. This comprehensively extracts features, reduces redundant features, avoids the adverse effects caused by mutual interference between different modalities, and effectively improves the accuracy of event prediction.
[0073] Based on the above embodiments, preferably, in an embodiment of the present invention, the generation of the multi-scale video features is as Figure 2 shown, and includes the following steps:
[0074] Step 210: Sample the target video stream with different sampling steps to obtain video frames of different sampling scales.
[0075] Among them, the sampling step can be understood as sampling video frames in the video stream with different steps. For example, the same video stream is sampled with steps of 1 frame, 3 frames, or 5 frames respectively to obtain three corresponding groups of video frames.
[0076] Step 220: Extract features from the video frames of different sampling scales to obtain multi-scale video features.
[0077] As described above, when predicting events for a video stream, it is necessary to determine the appearance features and action features in the video frame images in the video stream to predict upcoming events. In a video stream, since the appearance features change less over a period of time and the human action features change more within a certain period of time, video frames can be obtained with different sampling steps to extract appearance features and action features.
[0078] Let the length of the video frame sequence be p, then the video feature can be expressed as V ∈ R p×d . Where d represents the number of video frames in the target video stream. The multi-scale video processing module samples the video frame sequence feature V at different scales to generate video frame sequence features V1, V2,..., V n . Where n represents n sampling scales.
[0079] The method for multi-scale two-stream attention video-language event prediction provided by the present invention can effectively extract corresponding appearance features and action features by extracting video frames of different scales from the video stream with different sampling steps, which is convenient for model training and use, and improves the accuracy of event prediction.
[0080] Based on the above embodiments, preferably, in an embodiment of the present invention, the video frames of different sampling scales include: video frames of dense sampling scale, video frames of general sampling scale, and video frames of sparse sampling scale; correspondingly, extracting features from the video frames of different sampling scales to obtain multi-scale video features, as Figure 3 shown, includes the following steps:
[0081] Step 310, based on the video frames of the dense sampling scale and the pre-trained SlowFast model, obtain the first video feature of the video frames of the dense sampling scale.
[0082] Among them, the video frames of the dense sampling scale are video frames obtained based on a smaller sampling step. For example, the video frames sampled from the video stream with a sampling step of 1 can be used as the video frames of the dense sampling scale. The pre-trained SlowFast model is pre-trained on the Kinetics dataset and is used to extract the action features of video frames.
[0083] It can be understood that using the pre-trained SlowFast model is beneficial to reducing the overall training time of the event prediction model and improving the accuracy of event prediction.
[0084] Among them, d corresponding to the video frames of the dense sampling scale is 2304 frames.
[0085] It can be understood that the first video feature is that the action features of the video frames at the dense sampling scale can be extracted according to the video frames at the dense sampling scale and the pre-trained SlowFast model.
[0086] Step 320: Based on the video frames at the general sampling scale and the pre-trained ResNet-152 model, obtain the second video feature of the video frames at the general sampling scale.
[0087] Among them, the video frames at the general sampling scale are the video frames obtained based on a sampling step larger than the dense sampling scale. For example, the video frames obtained by sampling the video stream with a sampling step of 3 can be used as the video frames at the general sampling scale. The pre-trained ResNet-152 model is pre-trained on the ImageNet dataset and is used to extract the appearance features of video frames.
[0088] It can be understood that using the pre-trained ResNet-152 model is beneficial to reducing the overall training time of the event prediction model and improving the accuracy of event prediction.
[0089] It can be understood that the second video feature is that the appearance features of the video frames at the general sampling scale can be extracted according to the video frames at the general sampling scale and the pre-trained ResNet-152 model.
[0090] Among them, the number of frames d corresponding to the video frames at the general sampling scale is 2048 frames.
[0091] Step 330: Based on the video frames at the sparse sampling scale and the pre-trained SlowFast model, obtain the third video feature of the video frames at the sparse sampling scale; based on the video frames at the sparse sampling scale and the pre-trained ResNet-152 model, obtain the fourth video feature of the video frames at the sparse sampling scale; and splice the third video feature and the fourth video feature to obtain the fifth video feature.
[0092] It can be understood that video-language event prediction, as a video semantic understanding task, needs to extract various features of the video. ResNet-152 is excellent in extracting video appearance features, and SlowFast performs well in extracting video action features. The combination of the two can represent the video more comprehensively.
[0093] Among them, the video frames at the sparse sampling scale are the video frames obtained based on a sampling step larger than the general sampling scale. For example, the video frames obtained by sampling the video stream with a sampling step of 5 can be used as the video frames at the sparse sampling scale.
[0094] It can be understood that the third video feature is that the action features of the video frames at the sparse sampling scale can be extracted according to the video frames at the sparse sampling scale and the pre-trained SlowFast model. The fourth video feature is that the appearance features of the video frames at the sparse sampling scale can be extracted according to the video frames at the sparse sampling scale and the pre-trained ResNet-152 model.
[0095] Among them, the d corresponding to the video frames at the sparse sampling scale is 4352 (2304 + 2048) frames.
[0096] It can also be understood that the third video feature and the fourth video feature are concatenated to obtain the fifth video feature, so as to obtain the joint feature of the action feature and the appearance feature of the video frame, thus enriching the extracted video frame features.
[0097] Step 340: Determine the multi-scale video features based on the first video feature, the second video feature, and the fifth video feature.
[0098] It can be understood that the multi-scale features include the action features of the video frames, the appearance features, and the joint feature of the action feature and the appearance feature of the video frame, thus enriching the extracted video features.
[0099] The method for multi-scale two-stream attention video-language event prediction provided by the present invention extracts video frames at different sampling scales by using different feature extraction methods, so as to obtain multi-scale video features, extract video features in different directions, and enrich the video features.
[0100] Based on the above embodiments, preferably, in an embodiment of the present invention, the generation of the first fusion video features at different scales is as Figure 4 shown, and includes the following steps:
[0101] Step 410: Based on the single-modal feature conversion layer guided by future candidate events, fuse the video features at different scales in the multi-scale video features with the features of each future candidate event respectively to obtain the sixth video feature of the video features at different scales guided by future candidate events.
[0102] Among them, the single-modal feature conversion layer guided by future candidate events (which can be abbreviated as SEG / VEG) can be a layer of Transformer encoder.
[0103] Let the token-level length of the future candidate event be r. Then the feature of each future candidate event can be expressed as E i ∈R r×a, i ∈ {1, 2, …, a}. Here, a represents the number of future candidate events, and a corresponds to d in the previous text. It can be understood that a future candidate event is a text statement, and each of the future events can be converted into a sequence of words using existing techniques. Each word in the word sequence can correspond to a token. Therefore, the token-level length r is the length of the word sequence corresponding to each future candidate event.
[0104] In a possible implementation, first, the video features at each scale are concatenated with the features of each of the future candidate events to obtain a preliminary joint feature [V~E] ∈ R (p+r)×d , where ‘~’ represents the concatenation operation; then, the joint feature [V~E] ∈ R (p+r)×d is input into a single-modal feature transformation layer guided by the future candidate event. Using the self-attention mechanism, a representation V of the video V guided by the future candidate event E is generated E and a representation E of the future candidate event E guided by the video V V of the joint representation [V E ~E V ∈ R (p+r)×d . Finally, from the joint representation [V E ~E V , the representation V of the video V guided by the future candidate event E is split out E as the sixth video feature of the video features at different scales.
[0105] Among them, concatenating the video features at each scale with the features of each of the future candidate events can be understood as: concatenating in the dimensions of the video sequence length and the token-level length of the future candidate event, that is, connecting the features of a video frame with the features of a corresponding future candidate event. Specifically, the representation form of the joint feature [V~E] ∈ R (p+r)×d can be referred to. It can be understood that since a corresponds to d, both can represent the number of features of a certain modality. Therefore, R (p+r)×d is intended to mean that the feature concatenation is in the direction of the feature sequence lengths of different modalities, rather than the concatenation of the feature quantities of different modalities.
[0106] For ease of understanding, the sixth video feature is illustrated by an example. For example, if the first video feature of the video frames at the dense sampling scale is represented by V1, and the features of the future candidate events are E1 and E2, then the first video feature V1 is fused with E1 and E2 respectively to obtain and Combining the foregoing content, the fused features obtained based on the joint feature include V E and E VThe features of the two forms need to be split from the features of the two forms to obtain the representation V of the video V guided by the future candidate event E E , as the sixth video feature of the video features at each sampling scale. Therefore, the sixth video feature of the video frames at the dense sampling scale is determined as and and can be uniformly represented as V1 E .
[0107] Step 420: Based on the dual-stream video caption cross-modal fusion layer, fuse the video features of different scales in the multi-scale video features with the features of the captions corresponding to the target video stream respectively, and concatenate the fused features with the features of each future candidate event to obtain the video features of different scales guided by the captions; and input the video features of different scales guided by the captions into the single-modal feature conversion layer guided by the future candidate event to obtain the seventh video feature of the video features at each scale.
[0108] Among them, the dual-stream video caption cross-modal fusion layer (which can be abbreviated as SVVS) can be a single layer of Transformer encoder.
[0109] Let the token-level length of the entire caption sequence be q, then the features of the caption can be expressed as S ∈ R q×b . Among them, b represents the number of captions. It can be understood that the caption is a text sentence, and each caption can be converted into a sequence of words using existing technologies, and each word in the word sequence can correspond to a token. Therefore, the token-level length q is the length of the word sequence corresponding to each caption.
[0110] In a possible implementation, first concatenate the video features of each scale with the features of each caption to obtain the preliminary joint feature [S~V] ∈ R (q+r)×d , where '~' represents the concatenation operation; it can be understood that since b corresponds to d, both can represent the number of features of a certain modality, so R (q+r)×d is intended to express that the feature concatenation is the concatenation in the direction of the feature sequence length of different modalities, rather than the concatenation of the feature quantities of different modalities. Then, input the joint feature [S~V] ∈ R (q+r)×d into the single-modal feature conversion layer guided by the future candidate event, and use the self-attention mechanism to generate the representation V of the video V guided by the caption S S and the representation S of the caption S guided by the video V V of the joint representation [V S ~S V ∈ R (q+r)×d . Finally, from the joint representation [V S ~S VSplit out the representation V of the video V guided by the subtitle S S and the representation S of the subtitle S guided by the video V V , as the fusion feature of the video feature and subtitle feature at different scales. Among them, S V and V S The process of obtaining the features is similar to that of V in step 410 E . For the sake of brevity, it will not be elaborated here.
[0111] It can be understood that after obtaining V S , it is also necessary to concatenate with each of the future candidate events, and based on the concatenated features, obtain the video features V at different scales guided by the subtitle SE . Among them, if three sampling scales and 2 future candidate events are adopted, then V SE includes V1 SE , V2 SE and V3 SE ; V1 SE includes and V2 SE and V3 SE is similar to V1 SE , and will not be further described.
[0112] Step 430: Concatenate the sixth video feature and the seventh video feature corresponding to the video feature of each scale to obtain the first fusion video feature of each scale.
[0113] For the sake of understanding, combined with the previous example, the first fusion video feature is described. As mentioned above, if the sixth video feature at the dense sampling scale is V1 E , and the seventh video feature at the dense sampling scale is V1 SE , then the first fusion video feature at the dense sampling scale [V1 E ; V1 SE ∈ R p×2×d .
[0114] The method for multi-scale two-stream attention video-language event prediction provided by the present invention obtains the representation V of the video V guided by the future candidate event E by fusing the features of the future candidate event and the video features at different scales E , and fuses the representation V of the video V guided by the subtitle S S and the future candidate event E to obtain V SE , and combines V E and V SETogether as the first fused video feature, the associated features of the future candidate event, subtitle, and video frame are thus extracted, enabling the effective representation of the relationships among the three. Moreover, corresponding first fused video features exist for video frames at different sampling scales, enriching the diversity of video features.
[0115] Based on the above embodiments, preferably, in an embodiment of the present invention, the generation of the first fused subtitle features at different scales is as Figure 5 shown and includes the following steps:
[0116] Step 510: Based on the single-modal feature conversion layer guided by the future candidate event, fuse the features of the subtitle corresponding to the target video stream with the features of each future candidate event respectively to obtain the first subtitle feature guided by the future candidate event.
[0117] Among them, the formation process of the first subtitle feature guided by the future candidate event is similar to that of the video features at different scales guided by the future candidate event in step 410. For the sake of brevity, it will not be elaborated here. The finally formed first subtitle feature can be denoted as S E .
[0118] Step 520: Based on the dual-stream video subtitle cross-modal fusion layer, fuse the features of the subtitle corresponding to the target video stream with the multi-scale video features respectively to obtain the subtitle features guided by video frames at different scales; and based on the single-modal feature conversion layer guided by the future candidate event, fuse the fused features with the features of each future candidate event respectively to obtain multiple second subtitle features guided by the video.
[0119] Among them, the process of fusing the features of the subtitle corresponding to the target video stream with the video features at different scales can refer to step 420. After fusion, subtitle features S V guided by video frames at different scales are obtained. Then, using a method similar to step 420, S V is fused with the features of each future candidate event, thereby obtaining multiple second subtitle features S VE guided by the video. Among them, if three sampling scales and 2 future candidate events are adopted, S VE includes and includes and and and is similar and will not be further described.
[0120] Step 530: Concatenate the multiple first subtitle features and the multiple second subtitle features to obtain the first fused subtitle feature.
[0121] For ease of understanding, in combination with the previous example, the first fused subtitle feature is described. For example, the first subtitle feature guided by a future candidate event is S E , the subtitle feature guided by the video at the dense sampling scale and the future candidate event E are fused to obtain then the first fused subtitle feature at the dense sampling scale is
[0122] The method for multi-scale two-stream attention video-language event prediction provided by the present invention obtains the first subtitle feature S guided by a future candidate event by fusing the features of the future candidate event and the subtitle E , and the subtitle feature S guided by the video at different scales V and the future candidate event E are fused to obtain S VE , S E and S VE are jointly used as the first fused video feature, thereby extracting the correlation features of the future candidate event, the video frame and the subtitle, effectively representing the relationship among the three, and moreover, based on the video frames at different sampling scales having corresponding first fused subtitle features, the diversity of subtitle features is enriched.
[0123] Based on the above embodiments, preferably, in an embodiment of the present invention, the multi-scale two-stream attention video-language event prediction model further includes a subtitle and future candidate event feature extraction module. Correspondingly, the features of the subtitle and the features of the multiple future candidate events are generated based on the subtitle and future candidate event feature extraction module, including:
[0124] Input the subtitle corresponding to the target video stream into the subtitle and future candidate event feature extraction module to obtain the features of the subtitle;
[0125] Input the multiple future candidate events into the subtitle and future candidate event feature extraction module to obtain the features of the multiple future candidate events.
[0126] Among them, the subtitle and future candidate event feature extraction module can be, for example, a pre-trained RoBERTa-base model. The RoBERTa-base model is used to extract text features.
[0127] The method for multi-scale two-stream attention video-language event prediction provided by the present invention is beneficial to reducing the overall training time of the event prediction model and improving the accuracy of event prediction by using a pre-trained RoBERTa-base model.
[0128] Based on the above embodiments, preferably, in an embodiment of the present invention, the multi-scale two-stream attention video-language event prediction model further includes a multi-scale fusion module, which is used to fuse the first fusion video features of different scales to obtain second fusion video features, and is used to fuse the first fusion caption features of different scales to obtain second fusion caption features.
[0129] For ease of understanding, the second fusion video features and the second fusion caption features are illustrated by way of example.
[0130] If the first fusion video features at the dense sampling scale are [V1 E ; V1 SE ∈ R p×2×d and the first fusion video features at the general sampling scale are [V2 E ; V2 SE ∈ R p×2×d and the first fusion video features at the sparse sampling scale are [V3 E ; V3 SE ∈ R p ×2×d , then the corresponding second fusion video features are [V E ; V SE , where V E is obtained by summing up several matrices of V1 E , V2 E and V3 E , and V SE is obtained by summing up several matrices of V1 SE , V2 SE and V3 SE .
[0131] Similarly, if the first fusion caption features at the dense sampling scale are the first fusion caption features at the general sampling scale are the first fusion caption features at the sparse sampling scale are then the corresponding second fusion caption features are [S E ; S VE , where S VE is and obtained by summing up several matrices.
[0132] The method for multi-scale two-stream attention video-language event prediction provided by the present invention fuses the features of different scales by summing the video features and caption features of different scales respectively, so as to obtain the features after multi-scale fusion, which is convenient for subsequent processing.
[0133] Based on the above embodiments, preferably, in an embodiment of the present invention, obtaining the future candidate event prediction result of the target video stream based on the first fused video feature and the first fused subtitle feature is as follows Figure 6 shown, and includes the following steps:
[0134] Step 610: Compress the second fused video feature to obtain a compressed second fused video feature; and compress the second fused subtitle feature to obtain a compressed second fused subtitle feature.
[0135] It can be understood that compressing the second fused video feature and the second fused subtitle feature helps to reduce redundant features and also helps to speed up the prediction speed of the event prediction model.
[0136] Preferably, max pooling (MaxPool) can be used to compress the second fused video feature and the second fused subtitle feature.
[0137] Other methods, such as average pooling, can also be used for compression, and the present invention does not limit this.
[0138] Step 620: Perform event prediction based on the compressed second fused video feature to obtain multiple first scores corresponding to multiple future candidate events of the target video stream; and perform event prediction based on the compressed second fused subtitle feature to obtain multiple second scores corresponding to multiple future candidate events of the target video stream.
[0139] It can be understood that performing event prediction based on the second fused video feature and the second fused subtitle feature respectively is beneficial to distinguishing their prediction results and makes the model more flexible when performing event prediction.
[0140] Preferably, a multilayer perceptron (MLP) composed of two linear layers with the GELU function as the activation function can be used for event prediction.
[0141] Step 630: Add the first score of each future candidate event to the second score of each future candidate event to obtain the total score corresponding to each future candidate event of the target video stream.
[0142] It can be understood that there are multiple future candidate events, so there are corresponding scores for each future candidate event.
[0143] Step 640: Determine the future candidate events corresponding to the target video stream based on the total score corresponding to each future candidate event of the target video stream.
[0144] Preferably, the total score of each future candidate event can be normalized by SoftMax, and the future candidate event with the highest score is selected as the future candidate event corresponding to the target video stream.
[0145] The method for multi-scale two-stream attention video-language event prediction provided by the present invention performs event prediction based on the second fused video feature and the second fused caption feature respectively, which is beneficial to distinguishing the prediction results of the two and makes the model more flexible when performing event prediction. Moreover, summing the prediction results of the two is also beneficial to improving the accuracy of the model.
[0146] Based on the above embodiments, preferably, in an embodiment of the present invention, the multi-scale video feature is determined based on the first video feature, the second video feature, and the fifth video feature, as Figure 7 shown, and includes the following steps:
[0147] Step 710: Convert the first video feature, the second video feature, and the fifth video feature into the same dimension.
[0148] It can be understood that since the first video feature, the second video feature, and the fifth video feature are video features of different dimensions, they need to be converted into the same dimension for subsequent processing.
[0149] Specifically, a linear layer, such as a fully connected layer (FC), can be used to convert them into a unified dimension. The converted dimension can be 768 dimensions.
[0150] Step 720: Based on the Transformer encoder, perform temporal encoding on the first video feature, the second video feature, and the fifth video feature after dimension conversion respectively, to obtain the encoded first video feature, second video feature, and fifth video feature.
[0151] It can be understood that in order to extract the temporal features of the first video feature, the second video feature, and the fifth video feature, the temporal features of the first video feature, the second video feature, and the fifth video feature can be encoded based on one layer of the Transformer encoder. The principle of the Transformer encoder for performing temporal encoding on video features is to utilize the "attention" of the Self-Attention mechanism on the video frames of the target video stream.
[0152] Step 730: Use the encoded first video feature, second video feature, and fifth video feature as the multi-scale video feature.
[0153] It can be understood that the multi-scale video features obtained after encoding have temporal correlation, which is beneficial to improving the accuracy of event prediction.
[0154] The method for multi-scale two-stream attention video-language event prediction provided by the present invention determines multi-scale features by performing dimensional transformation and encoding on video features of different scales, so that the obtained multi-scale features have temporality, which is convenient for subsequent processing and is beneficial to improving the accuracy of event prediction.
[0155] Figure 8 It is a schematic framework diagram of the method for multi-scale two-stream attention video-language event prediction provided by the present invention.
[0156] As Figure 8 shown, the framework of the method for multi-scale two-stream attention video-language event prediction provided by the present invention includes several parts: input representation, multi-scale sampling and encoding, cross-modal fusion V1, cross-modal fusion V2, cross-modal fusion V3, multi-scale fusion, and prediction output.
[0157] Among them, for input representation, different models are respectively used to extract corresponding features from future events, captions, and videos. Among them, the video can be a target video stream, and the future event is a future candidate event preset according to the target video stream, which can be in the form of a text. For example, it can be "The women in the white shirt...". The caption is the corresponding caption in the target video stream, which can be a text. For example, it can be "Oh yeah!Maybe a shake...". The future event and the caption are respectively input into the RoBERTa-base model to extract features to obtain the feature E of the future event and the feature S of the caption respectively. Then, multi-scale video features can be generated based on the slowfast model, the ResNet-152 model, and the multi-scale sampler. Specifically, the method for generating multi-scale video features described above can be referred to to extract the multi-scale video features in the target video stream.
[0158] Secondly, the multi-scale video features obtained by the multi-scale sampler are respectively input into the corresponding fully connected layer FC and a layer of Transformer encoder (corresponding to T-E in the figure) to obtain the encoded multi-scale video features V1, V2, and V3.
[0159] Then, cross-modal fusion is respectively performed on V1, V2, and V3 to obtain the first fused video features and the first fused caption features after fusion at different scales. Figure 8 Specifically shows the acquisition process of the first fused video features and the first fused caption features at one scale. Finally, the first fused caption feature corresponding to the feature V1 at one scale is First fused video feature [V1E ; V1 SE . Specifically, the first fused video feature and the first fused subtitle feature can refer to the relevant descriptions in the previous text and will not be elaborated here.
[0160] It can be understood that Figure 8 illustrates in detail the process of obtaining the first fused video feature and the first fused subtitle feature corresponding to a feature V1 of one scale. The processes of obtaining the first fused video feature and the first fused subtitle feature of V2 and V3 are similar to that of V1 and will not be elaborated here.
[0161] Finally, after cross-modal fusion to obtain the first fused video features and the first fused subtitle features corresponding to V1, V2, and V3, the first fused video features of different scales are fused, and the first fused subtitle features of different scales are fused to obtain the second fused video feature [V E ; V SE and the second fused subtitle feature [S E ; S VE . The second fused video feature [V E ; V SE and the second fused subtitle feature [S E ; S VE are respectively input into MaxPool to compress the second fused video feature and the second fused subtitle feature. The compressed features are respectively input into MLP for event prediction. Based on the second fused video feature [V E ; V SE , multiple first scores corresponding to multiple future candidate events of the target video stream are obtained, and based on the second fused subtitle feature [S E ; S VE , multiple second scores corresponding to multiple future candidate events of the target video stream are obtained. Then, the multiple first scores and the second scores of each future candidate event are added to obtain the total score of each future candidate event. The total scores of each future candidate event are normalized through SoftMax, and the future candidate event with the highest score is selected as the future candidate event corresponding to the target video stream.
[0162] The following combines Figure 8 and the experimental results in Table 1 to illustrate the achievable effects of the present invention.
[0163] Table 1
[0164] Model Accuracy (%) Backbone 67.33 Backbone + Multi-scale Sampling 68.08 Backbone + Cross-modal Fusion 68.62 Full Model 69.65
[0165] As shown in Table 1, among them, the backbone model is Figure 8The model obtained by removing the multi-scale sampler and multi-scale fusion, and adapting other parts after removing cross-modal fusion. Based on the backbone model, the accuracy rate is 67.33%. The model corresponding to the backbone + multi-scale sampling is the model after adapting other parts after removing the cross-modal fusion in Figure 8 , and its corresponding accuracy rate is 68.08%. The model corresponding to the backbone + cross-modal fusion is the model after adapting other parts after removing the multi-scale sampler and multi-scale fusion in Figure 8 , and its corresponding accuracy rate is 68.62%. Finally, the accuracy rate obtained by using the complete model in Figure 8 is 69.65%.
[0166] Therefore, it can be known that the method for predicting multi-scale two-stream attention video language events provided by the present invention can effectively improve the accuracy of event prediction.
[0167] Figure 9 is a schematic diagram of the device for predicting multi-scale two-stream attention video language events provided by the present invention. As shown in Figure 9 , the device for predicting multi-scale two-stream attention video language events provided by the embodiments of the present invention includes:
[0168] An acquisition module 910, configured to acquire original input data; wherein, the original input data includes a target video stream, subtitles corresponding to the target video stream, and a plurality of future candidate events;
[0169] A processing module 920, configured to input the original input data into a multi-scale two-stream attention video language event prediction model to obtain an event prediction result of the target video stream;
[0170] Wherein, the multi-scale two-stream attention video language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module;
[0171] The multi-scale video processing module is configured to generate multi-scale video features based on video frames in the target video stream; the two-stream cross-modal fusion module is configured to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitles, the features of the plurality of future candidate events, and the multi-scale video features; the event prediction module is configured to respectively obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales, and determine a final event prediction result of the target video stream based on the event prediction results.
[0172] The device for multi-scale two-stream attention video-language event prediction provided by the embodiments of the present invention specifically executes the process of the method embodiment of the above multi-scale two-stream attention video-language event prediction. For the specific content, please refer to the content of the method embodiment of the above multi-scale two-stream attention video-language event prediction, which will not be elaborated here.
[0173] The device for multi-scale two-stream attention video-language event prediction provided by the present invention obtains multi-scale video features by performing multi-scale processing on the video, making the extracted video features more reasonable. Based on the multi-scale video features, the features of the subtitle, and the features of multiple future candidate events, it fuses to obtain the first fusion video features and the first fusion subtitle features at different scales, and respectively performs event prediction based on the first fusion video features and the first fusion subtitle features at different scales, and then combines the prediction results to determine the final event prediction result, comprehensively extracting features, reducing redundant features, avoiding the adverse effects caused by mutual interference between different modalities, and effectively improving the accuracy of event prediction.
[0174] Figure 10 An example of the physical structure diagram of an electronic device is as Figure 10 shown. The electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the method for multi-scale two-stream attention video-language event prediction, including: obtaining the original input data; where the original input data includes a target video stream, the subtitle corresponding to the target video stream, and multiple future candidate events; inputting the original input data into the multi-scale two-stream attention video-language event prediction model to obtain the event prediction result of the target video stream; where the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module; the multi-scale video processing module is used to generate multi-scale video features based on the video frames in the target video stream; the two-stream cross-modal fusion module is used to generate the first fusion video features and the first fusion subtitle features at different scales based on the features of the subtitle, the features of the multiple future candidate events, and the multi-scale video features; the event prediction module is used to respectively obtain event prediction results based on the first fusion video features and the first fusion subtitle features at different scales, and determine the final event prediction result of the target video stream based on the event prediction results.
[0175] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0176] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for multi-scale two-stream attention video-language event prediction provided by the present invention, including: obtaining original input data; wherein, the original input data includes a target video stream, the subtitles corresponding to the target video stream, and multiple future candidate events; inputting the original input data into a multi-scale two-stream attention video-language event prediction model to obtain an event prediction result of the target video stream; wherein, the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module; the multi-scale video processing module is used to generate multi-scale video features based on the video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitles, the features of the multiple future candidate events, and the multi-scale video features; the event prediction module is used to obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales respectively, and determine the final event prediction result of the target video stream based on the event prediction results.
[0177] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the method for multi-scale two-stream attention video-language event prediction provided by the present invention, including: obtaining original input data; wherein, the original input data includes a target video stream, the subtitle corresponding to the target video stream, and a plurality of future candidate events; inputting the original input data into a multi-scale two-stream attention video-language event prediction model to obtain an event prediction result of the target video stream; wherein, the multi-scale two-stream attention video-language event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module; the multi-scale video processing module is used to generate multi-scale video features based on the video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitle, the features of the plurality of future candidate events, and the multi-scale video features; the event prediction module is used to respectively obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales, and determine the final event prediction result of the target video stream based on the event prediction results.
[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for multi-scale two-stream attention video-language event prediction, characterized in that, Including: Obtain original input data; wherein, the original input data includes a target video stream, subtitles corresponding to the target video stream, and multiple future candidate events; Input the original input data into a multi-scale dual-stream attention video-language event prediction model to obtain an event prediction result of the target video stream; Wherein, the multi-scale dual-stream attention video-language event prediction model includes a multi-scale video processing module, a dual-stream cross-modal fusion module, and an event prediction module; The multi-scale video processing module is used to generate multi-scale video features based on video frames in the target video stream; the dual-stream cross-modal fusion module is used to generate first fusion video features at different scales and first fusion subtitle features at different scales based on the features of the subtitles, the features of the multiple future candidate events, and the multi-scale video features; the event prediction module is used to obtain event prediction results based on the first fusion video features at different scales and the first fusion subtitle features at different scales, and determine a final event prediction result of the target video stream based on the event prediction results; Wherein, the generation of the first fusion video features at different scales includes the following steps: Based on a single-modal feature transformation layer guided by future candidate events, fuse video features at different scales in the multi-scale video features with the features of each future candidate event respectively to obtain sixth video features of the video features at different scales guided by future candidate events; Based on a dual-stream video-subtitle cross-modal fusion layer, fuse video features at different scales in the multi-scale video features with the features of the subtitles corresponding to the target video stream respectively, and concatenate the fused features with the features of each future candidate event to obtain video features at different scales guided by subtitles; and input the video features at different scales guided by subtitles into a single-modal feature transformation layer guided by future candidate events to obtain seventh video features of the video features at each scale; Concatenate the sixth video features and the seventh video features corresponding to the video features at each scale to obtain first fusion video features at each scale; The generation of the first fusion subtitle features at different scales includes the following steps: Based on a single-modal feature transformation layer guided by future candidate events, fuse the features of the subtitles corresponding to the target video stream with the features of each future candidate event respectively to obtain first subtitle features guided by future candidate events; Based on a dual-stream video-subtitle cross-modal fusion layer, fuse the features of the subtitles corresponding to the target video stream with the multi-scale video features respectively to obtain subtitle features guided by video frames at different scales; and based on a single-modal feature transformation layer guided by future candidate events, fuse the fused features with the features of each future candidate event respectively to obtain multiple second subtitle features guided by video; Concatenate the multiple first subtitle features and the multiple second subtitle features to obtain first fusion subtitle features.
2. The method for multi-scale two-stream attention video-language event prediction according to claim 1, wherein The generation of the multi-scale video features includes: Sample the target video stream with different sampling steps to obtain video frames at different sampling scales; Extract features from the video frames at different sampling scales to obtain multi-scale video features.
3. The method for multi-scale two-stream attention video-linguistic event prediction according to claim 2, wherein The video frames of different sampling scales include: video frames of a dense sampling scale, video frames of a general sampling scale, and video frames of a sparse sampling scale; correspondingly, extracting features from the video frames of different sampling scales to obtain multi-scale video features includes: Based on the video frames of the dense sampling scale and a pre-trained SlowFast model, obtaining first video features of the video frames of the dense sampling scale; Based on the video frames of the general sampling scale and a pre-trained ResNet-152 model, obtaining second video features of the video frames of the general sampling scale; Based on the video frames of the sparse sampling scale and a pre-trained SlowFast model, obtaining third video features of the video frames of the sparse sampling scale; based on the video frames of the sparse sampling scale and a pre-trained ResNet-152 model, obtaining fourth video features of the video frames of the sparse sampling scale; and splicing the third video features and the fourth video features to obtain fifth video features; Determining multi-scale video features based on the first video features, the second video features, and the fifth video features.
4. The method for multi-scale two-stream attention video-linguistic event prediction according to claim 1, wherein The multi-scale two-stream attention video-language event prediction model further includes a caption and future candidate event feature extraction module. Correspondingly, the features of the caption and the features of the multiple future candidate events are generated based on the caption and future candidate event feature extraction module, including: Inputting the caption corresponding to the target video stream into the caption and future candidate event feature extraction module to obtain the features of the caption; Inputting the multiple future candidate events into the caption and future candidate event feature extraction module to obtain the features of the multiple future candidate events.
5. The method for multi-scale two-stream attention video-linguistic event prediction according to claim 1, wherein The multi-scale two-stream attention video-language event prediction model further includes a multi-scale fusion module. The multi-scale fusion module is used to fuse the first fusion video features of different scales to obtain second fusion video features, and is used to fuse the first fusion caption features of different scales to obtain second fusion caption features.
6. The method for multi-scale two-stream attention video-language event prediction according to claim 5, characterized in that Obtaining the future candidate event prediction result of the target video stream based on the first fusion video features and the first fusion caption features includes: Compressing the second fusion video features to obtain compressed second fusion video features; and compressing the second fusion caption features to obtain compressed second fusion caption features; Performing event prediction based on the compressed second fusion video features to obtain multiple first scores of multiple future candidate events corresponding to the target video stream; and performing event prediction based on the compressed second fusion caption features to obtain multiple second scores of multiple future candidate events corresponding to the target video stream; Adding the first score of each future candidate event to the second score of each future candidate event to obtain the total score of each future candidate event corresponding to the target video stream; Determining the future candidate events corresponding to the target video stream based on the total scores of each future candidate event corresponding to the target video stream.
7. The method for multi-scale two-stream attention video-language event prediction according to claim 3, wherein, Determining the multi-scale video feature based on the first video feature, the second video feature, and the fifth video feature includes: Converting the first video feature, the second video feature, and the fifth video feature into the same dimension; Based on the Transformer encoder, performing temporal encoding on the first video feature, the second video feature, and the fifth video feature after dimension conversion respectively to obtain the encoded first video feature, second video feature, and fifth video feature; Using the encoded first video feature, second video feature, and fifth video feature as the multi-scale video feature.
8. An apparatus for multi-scale two-stream attention video-language event prediction, characterized in that, Including: An acquisition module for acquiring original input data; wherein, the original input data includes a target video stream, the subtitle corresponding to the target video stream, and multiple future candidate events; A processing module for inputting the original input data into a multi-scale two-stream attention video-linguistic event prediction model to obtain the event prediction result of the target video stream; Wherein, the multi-scale two-stream attention video-linguistic event prediction model includes a multi-scale video processing module, a two-stream cross-modal fusion module, and an event prediction module; The multi-scale video processing module is used to generate multi-scale video features based on the video frames in the target video stream; the two-stream cross-modal fusion module is used to generate first fusion video features of different scales and first fusion subtitle features of different scales based on the features of the subtitle, the features of the multiple future candidate events, and the multi-scale video features; the event prediction module is used to obtain event prediction results based on the first fusion video features of different scales and the first fusion subtitle features of different scales respectively, and determine the final event prediction result of the target video stream based on the event prediction results; Wherein, the generation of the first fusion video features of different scales includes the following steps: Based on the future candidate event-guided single-modal feature conversion layer, respectively fusing the video features of different scales in the multi-scale video features with the features of each future candidate event to obtain the sixth video feature of the video features of different scales guided by the future candidate event; Based on the two-stream video-subtitle cross-modal fusion layer, respectively fusing the video features of different scales in the multi-scale video features with the features of the subtitle corresponding to the target video stream, and concatenating the fused features with the features of each future candidate event to obtain the video features of different scales guided by the subtitle; and inputting the video features of different scales guided by the subtitle into the future candidate event-guided single-modal feature conversion layer to obtain the seventh video feature of the video features of each scale; Concatenating the sixth video feature and the seventh video feature corresponding to the video features of each scale to obtain the first fusion video feature of each scale; The generation of the first fusion subtitle features of different scales includes the following steps: Based on the future candidate event-guided single-modal feature conversion layer, respectively fusing the features of the subtitle corresponding to the target video stream with the features of each future candidate event to obtain the first subtitle feature guided by the future candidate event; Based on the cross-modal fusion layer of dual-stream video subtitles, the features of the subtitles corresponding to the target video stream are respectively fused with the multi-scale video features to obtain subtitle features guided by video frames of different scales; and based on the single-modal feature conversion layer guided by future candidate events, the fused features are respectively fused with the features of each future candidate event to obtain multiple second subtitle features guided by the video. The multiple first subtitle features and the multiple second subtitle features are concatenated to obtain the first fused subtitle feature.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for multi-scale dual-stream attention video-language event prediction according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for multi-scale dual-stream attention video-language event prediction according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for multi-scale dual-stream attention video-language event prediction according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target behavior perception method for audio and video cross-modal feature expression
CN117011763A