Multi-modal video fusion method and device for wide-area visual internet, equipment and medium
Through synchronous shooting and multimodal video data fusion methods, the problem of insufficient video data fusion accuracy in wide-area visual networking systems was solved, high-precision reconstructed videos were generated, and stable monitoring perception of large areas was achieved.
Patent Information
- Application Number
- CN202511007573.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The existing wide-area visual network system is unable to effectively integrate video data from remote sensing satellites, drones, and ground towers, resulting in insufficient detection accuracy for targets in large areas and making it difficult to achieve high-precision monitoring and information services.
Multimodal video data is captured synchronously by cameras of remote sensing satellites, drones, and ground towers. Feature extraction and grouping processing are performed using the multimodal shallow feature extraction module. Parallel and global cyclic feature refinement are performed in combination with the fusion feature loop refinement module. Finally, upsampling is performed in the video reconstruction module to generate high-resolution and wide-field-of-view reconstructed videos.
It achieves high-precision detection of targets in a large area, provides stable and flexible intelligent monitoring perception, integrates video data of different resolutions and fields of view, and generates reconstructed videos with a wide field of view and high resolution.
Smart Images

Figure CN120510485B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a multimodal video fusion method, device, equipment and medium for wide-area visual networking. Background Art
[0002] Wide-area video networks (WAVs) have observation ranges of up to tens of kilometers. However, effective communication and information services are still lacking for scenarios with observation ranges of hundreds or even thousands of kilometers (such as remote suburbs, offshore areas, and border regions). A space-ground wide-area video network leverages the advantages of high-rise towers, drones, and satellites to provide high-definition video surveillance and information services within observation ranges of tens to thousands of kilometers.
[0003] Currently, the integrated space-ground wide-area visual network includes a variety of video acquisition devices, including remote sensing satellites, drones, and ground towers. The video data collected by different video acquisition devices has different application scenarios and characteristics: remote sensing satellites provide a large field of view but low resolution, making them suitable for large-scale monitoring, but with limited accuracy; drones and ground towers have a smaller field of view and higher resolution, making them suitable for detailed local observations, but with a narrower monitoring range; drones also offer greater flexibility and can be deployed on demand to areas that towers cannot reach, with low deployment costs, enabling more refined monitoring of larger areas than towers. However, none of these video data can achieve high-precision detection of targets within a large area. Summary of the Invention
[0004] Based on the above technical problems, the present invention provides a multimodal video fusion method, device, equipment and medium for wide-area visual networking, aiming to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of the present invention provides a multimodal video fusion method for wide-area visual networking, the method comprising:
[0006] The target scene is captured synchronously by cameras of remote sensing satellites, drones, and ground towers to obtain multimodal video data, wherein the multimodal video data includes: satellite video data, drone video data, and ground tower video data;
[0007] Inputting the multimodal video data into a multimodal shallow feature extraction module to obtain multimodal fusion shallow features;
[0008] The multimodal fusion shallow features are grouped and processed in chronological order to obtain multiple groups of multimodal fusion shallow features, each group of multimodal fusion shallow features includes N multimodal fusion shallow features that are continuous in time, where N is an integer greater than 1;
[0009] input the multi-group multi-modal refined features into a video reconstruction module, up-sample the multi-group multi-modal refined features, and obtain a reconstructed video, wherein the reconstructed video fuses information provided by the satellite video data, the unmanned aerial vehicle video data and the ground tower video data.
[0010] input the multi-group multi-modal refined features into a video reconstruction module, up-sample the multi-group multi-modal refined features, and obtain a reconstructed video, wherein the reconstructed video fuses information provided by the satellite video data, the unmanned aerial vehicle video data and the ground tower video data.
[0011] The second aspect of the present application provides a multi-modal video fusion device for a wide-area visual internet, and the device comprises:
[0012] The data acquisition module is configured to synchronously capture a target scene by a camera of a remote sensing satellite, an unmanned aerial vehicle and a ground tower, and obtain multi-modal video data, wherein the multi-modal video data comprises satellite video data, unmanned aerial vehicle video data and ground tower video data.
[0013] The feature extraction module is configured to input the multi-modal video data into a multi-modal shallow feature extraction module, and obtain multi-modal fusion shallow features.
[0014] The feature grouping module is configured to group the multi-modal fusion shallow features in a time sequence, and obtain multi-group multi-modal fusion shallow features, wherein each group of multi-modal fusion shallow features comprises N multi-modal fusion shallow features that are continuous in time, and N is an integer greater than 1.
[0015] The feature refinement module is configured to input the multi-group multi-modal fusion shallow features into a fusion feature cyclic refinement module, perform parallel feature refinement in each group by the fusion feature cyclic refinement module, perform global cyclic feature refinement between different groups by the fusion feature cyclic refinement module, and obtain multi-group multi-modal refined features.
[0016] The video reconstruction module is configured to input the multi-group multi-modal refined features into a video reconstruction module, up-sample the multi-group multi-modal refined features, and obtain a reconstructed video, wherein the reconstructed video fuses information provided by the satellite video data, the unmanned aerial vehicle video data and the ground tower video data.
[0017] The third aspect of the present application provides an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program is executed by the processor to implement the multi-modal video fusion method for a wide-area visual internet according to the first aspect of the present application.
[0018] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multimodal video fusion method for wide-area visual networking as described in the first aspect of the present invention is implemented.
[0019] In the multimodal video fusion method for wide-area visual networking proposed in the present invention, the cameras of remote sensing satellites, drones and ground towers are used to synchronously shoot the target scene, and multimodal video data with different resolutions and fields of view are obtained respectively, and the multimodal video data are cross-domain fused: first, shallow features are extracted from the multimodal video data to obtain multimodal fusion shallow features and group them to obtain multiple groups of multimodal fusion shallow features; then, the multiple groups of multimodal fusion shallow features are input into the fusion feature loop refinement module, and the fusion feature loop refinement module combines the global loop structure and the local loop structure. The method combines partial parallel processing with modeling of the temporal dependency of the video, thereby performing parallel feature refinement within each group of multiple multimodal fusion shallow features, and performing global cyclic feature refinement between different groups of multiple multimodal fusion shallow features. Thus, the temporal dependency of the video frames is utilized to gradually refine the multimodal fusion shallow features, resulting in multiple groups of multimodal refined features. Finally, the multiple groups of multimodal refined features are input into a video reconstruction module to obtain a reconstructed video that integrates information provided by satellite video data, drone video data, and ground tower video data. In this way, the present invention comprehensively utilizes multimodal video data of different resolutions and fields of view, and intelligently fuses the multimodal video data through a multimodal shallow feature extraction module, a fusion feature cyclic refinement module, and a video reconstruction module, generating a reconstructed video with a wide field of view and high resolution. This achieves high-precision detection of targets within a large area, and realizes stable and flexible intelligent monitoring perception in a wide-area visual network. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 This is a flowchart of a multimodal video fusion method for wide-area visual networking according to an embodiment of the present invention;
[0022] Figure 2 is a flowchart showing a multimodal video fusion method for wide-area visual networking according to another embodiment of the present invention;
[0023] Figure 3This is a structural block diagram of a multimodal video fusion device for wide-area visual networking provided by one embodiment of the present invention;
[0024] Figure 4 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] Please refer to Figure 1 , Figure 1 FIG. 1 is a flowchart showing a multimodal video fusion method for wide area visual networking according to an embodiment of the present invention. Figure 1 As shown, the multimodal video fusion method for wide-area visual networking provided by this embodiment includes at least the following steps:
[0027] Step S11: synchronously capture the target scene through the cameras of the remote sensing satellite, the UAV, and the ground tower to obtain multimodal video data.
[0028] In this embodiment, for the integrated space-ground wide-area video network, remote sensing satellites, drones, and ground towers can simultaneously capture the same target scene through their cameras, generating multimodal video data for the same target scene. The multimodal video data includes satellite video data, drone video data, and ground tower video data.
[0029] Remote sensing satellites offer a wide range and macroscopic coverage; drones offer flexibility and maneuverability; and ground-based towers offer high resolution and strong real-time performance. Remote sensing video data, with its wide observation range and strong macroscopic coverage, is suitable for applications such as large-scale landform analysis and land use classification. However, remote sensing platforms are limited by spatial resolution, resulting in insufficient perception accuracy for small-scale targets. Furthermore, their temporal resolution is typically several hours or even days, making it difficult to achieve full-time continuous coverage. Drones offer flexibility and maneuverability, enabling rapid response and localized high-resolution observations. However, due to limitations in battery capacity and number, drones typically have a flight range of only tens of kilometers and a flight time of approximately half an hour, limiting their spatial and temporal scalability. Ground-based towers offer high spatial resolution and strong real-time performance, enabling them to capture information on small objects, subtle landform features, and rapidly changing targets. However, their limited observation range and single perspective make them unsuitable for wide-area, complex scenarios.
[0030] It can be understood that the multimodal video data of this embodiment is video data with different fields of view and resolutions, wherein the first field of view corresponding to the satellite video data is larger than the second field of view corresponding to the drone video data, the first field of view corresponding to the satellite video data is larger than the third field of view corresponding to the ground tower video data, and the first resolution corresponding to the satellite video data is smaller than the second resolution corresponding to the drone video data, and the first resolution corresponding to the satellite video data is smaller than the third resolution corresponding to the ground tower video data.
[0031] Step S12: inputting the multimodal video data into a multimodal shallow feature extraction module to obtain multimodal fusion shallow features.
[0032] In this embodiment, a multimodal video fusion network is pre-trained and includes at least a multimodal shallow feature extraction module, a fusion feature recursive refinement module, and a video reconstruction module. Multimodal video data can be input into the multimodal shallow feature extraction module, which then extracts shallow features from multiple modal video frames within the multimodal video data and performs a preliminary fusion of the multiple modalities to produce multimodal fused shallow features. In this embodiment, shallow features are low-level visual information, primarily including basic structural features such as edges, texture, and color.
[0033] In an optional embodiment, the multimodal shallow feature extraction module can use a convolution kernel to extract spatial features from the input multimodal video data, and then use multiple moving window-based multi-head attention feature extraction submodules to extract shallow features.
[0034] Step S13: grouping the multimodal fusion shallow features in chronological order to obtain multiple groups of multimodal fusion shallow features.
[0035] In this embodiment, the multimodal fusion shallow features output by the multimodal shallow feature extraction module are multiple multimodal fusion shallow features. The multiple multimodal fusion shallow features correspond to the time dependency of the video frames. It can be understood that each multimodal fusion shallow feature corresponds to one video frame. In this embodiment, the multiple multimodal fusion shallow features can be grouped and processed according to the time sequence of the video frames to obtain multiple groups of multimodal fusion shallow features. Each group of multimodal fusion shallow features includes N multimodal fusion shallow features that are continuous in time. That is, each group of multimodal fusion shallow features includes: N multimodal fusion shallow features corresponding to N adjacent video frames that are continuous in time, where N is an integer greater than 1.
[0036] In an optional embodiment, the specific value of N can be freely determined based on computing resources, storage resources, or computing power. For example, when computing resources are sufficient, N is 8; when computing resources are scarce, N is 4. This embodiment does not impose any restrictions on this.
[0037] Step S14: Input the multiple groups of multimodal fused shallow features into the fusion feature cyclic refinement module, perform parallel feature refinement in each group through the fusion feature cyclic refinement module, and perform global cyclic feature refinement between different groups through the fusion feature cyclic refinement module to obtain multiple groups of multimodal refined features.
[0038] In this embodiment, after obtaining multiple groups of multimodal fusion shallow features, the multiple groups of multimodal fusion shallow features can be input into the fusion feature loop refinement module. The fusion feature loop refinement module of this embodiment adopts a structure that combines a global loop structure and local parallel processing, and uses the time dependency of the video frame to gradually refine the multimodal fusion shallow features. Among them, for multiple groups of multimodal fusion shallow features, the fusion feature loop refinement module performs parallel feature refinement in each group of the multiple groups of multimodal fusion shallow features, that is, the fusion feature loop refinement module performs feature refinement processing on the N multimodal fusion shallow features in each group of multimodal fusion shallow features at the same time.
[0039] In addition, for multiple groups of multimodal fusion shallow features, the fusion feature loop refinement module performs global loop feature refinement between different groups in the multiple groups of multimodal fusion shallow features, so as to refine the multimodal fusion shallow features of different groups in sequence based on the time sequence, thereby realizing global loop feature refinement.
[0040] Based on this, the fusion feature cyclic refinement module of this embodiment performs parallel feature refinement within each group of multiple groups of multimodal fusion shallow features, and performs global cyclic feature refinement between different groups of multiple groups of multimodal fusion shallow features, to obtain multiple groups of multimodal refined features corresponding to the multiple groups of multimodal fusion shallow features output by the fusion feature cyclic refinement module.
[0041] Step S15: inputting the multiple sets of multimodal refined features into a video reconstruction module, upsampling the multiple sets of multimodal refined features, and obtaining a reconstructed video.
[0042] In this embodiment, after obtaining multiple sets of multi-modal refined features, the multiple sets of multi-modal refined features can be input into a video reconstruction module, the multiple sets of multi-modal refined features are up-sampled by the video reconstruction module, and a reconstructed video is output, which fuses information provided by the satellite video data, the unmanned aerial vehicle video data, and the ground tower video data, and is a reconstructed video with a wide field of view and high resolution. In an optional example, the video reconstruction module at least includes a pixel shuffle up-sampling module, and the multiple sets of multi-modal refined features can be up-sampled by the pixel shuffle up-sampling module, and the reconstructed video is output.
[0043] In this embodiment, the multi-modal video data with different resolutions and fields of view are comprehensively utilized, the multi-modal video data is intelligently fused by deep learning, a reconstructed video with a wide field of view and high resolution is generated, high-precision detection of targets in a large range is achieved, and stable and flexible intelligent monitoring and sensing of the wide-area visual internet is achieved. In this embodiment, the global cyclic structure and the local parallel processing are combined by the fusion feature cyclic refinement module, the time-dependent relationship of the video is modeled, parallel feature refinement is performed in each set of multi-modal fusion shallow features, and global cyclic feature refinement is performed between different sets of multi-modal fusion shallow features, so that the multi-modal fusion shallow features are gradually refined by using the time-dependent relationship of the video frames, the balance between the model parameter size and the modeling of long time-dependent relationships is achieved, and a certain parallelism is also achieved.
[0044] In combination with the above embodiments, in an implementation, the present application further provides a multi-modal video fusion method for a wide-area visual internet, in which the step S12 can specifically include steps S21 and S25:
[0045] Step S21: inputting the satellite video data into a first feature extraction submodule in the multi-modal shallow feature extraction module to obtain convolutional satellite video features.
[0046] In this embodiment, for the input satellite video data, unmanned aerial vehicle video data, and ground tower video data from different modalities of remote sensing satellites, unmanned aerial vehicles, and ground towers, different feature extraction submodules are used to perform preliminary feature extraction on video frames in the video data of different modalities. The first feature extraction submodule in the multi-modal shallow feature extraction module is a feature extraction submodule corresponding to the satellite video data. The satellite video data can be input into the first feature extraction submodule in the multi-modal shallow feature extraction module to obtain convolutional satellite video features output by the first feature extraction submodule.
[0047] Step S22: inputting the unmanned aerial vehicle video data into a second feature extraction submodule in the multi-modal shallow feature extraction module to obtain convolutional unmanned aerial vehicle video features.
[0048] In this embodiment, the second feature extraction submodule in the multimodal shallow feature extraction module is a feature extraction submodule corresponding to the drone video data. The drone video data can be input into the second feature extraction submodule in the multimodal shallow feature extraction module to obtain the convolved drone video features output by the second feature extraction submodule.
[0049] Step S23: inputting the ground tower video data into the third feature extraction submodule in the multimodal shallow feature extraction module to obtain the convolved ground tower video features.
[0050] In this embodiment, the third feature extraction submodule in the multimodal shallow feature extraction module is a feature extraction submodule corresponding to the ground tower video data. The ground tower video data can be input into the third feature extraction submodule in the multimodal shallow feature extraction module to obtain the convolved ground tower video features output by the third feature extraction submodule.
[0051] In a specific example, for a video frame in the input video data , the corresponding feature extraction submodule is used for preliminary feature extraction:
[0052] ;
[0053] in, is the convolutional video feature corresponding to different modalities, are feature extraction submodules corresponding to different modalities, where T is the number of video frames, m represents the modality, Represent satellite, drone, and ground tower modes respectively.
[0054] Since the resolutions of satellite video data, drone video data, and ground tower video data are generally different, this embodiment configures different convolution kernel parameters for the first feature extraction submodule, the second feature extraction submodule, and the third feature extraction submodule to ensure that the extracted feature dimensions remain consistent (the feature dimensions of the convolved satellite video features, the convolved drone video features, and the convolved ground tower video features are consistent). The first parameter value of the convolution kernel in the first feature extraction submodule is adapted to the video resolution of the satellite video data, the second parameter value of the convolution kernel in the second feature extraction submodule is adapted to the video resolution of the drone video data, and the third parameter value of the convolution kernel in the third feature extraction submodule is adapted to the video resolution of the ground tower video data.
[0055] Furthermore, in an optional embodiment, the first feature extraction submodule includes a first number of first convolution kernels; the first feature extraction submodule includes a second number of second convolution kernels; and the third feature extraction submodule includes a third number of third convolution kernels. The second number and the third number are each greater than the first number, and the step size of the second convolution kernel and the step size of the third convolution kernel are each greater than the step size of the first convolution kernel.
[0056] Step S24: The convolved satellite video features, the convolved drone video features, and the convolved ground tower video features are respectively input into a plurality of moving window-based multi-head attention feature extraction submodules to obtain satellite video shallow features, drone video shallow features, and ground tower video shallow features.
[0057] In this embodiment, the multimodal shallow feature extraction module also includes a plurality of moving window-based multi-head attention feature extraction submodules (such as a plurality of moving window-based Swin Transformer blocks). For the convolutional satellite video features, the convolutional drone video features, and the convolutional ground tower video features, the convolutional satellite video features, the convolutional drone video features, and the convolutional ground tower video features can be respectively input into a plurality of moving window-based multi-head attention feature extraction submodules, thereby obtaining the satellite video shallow features, drone video shallow features, and ground tower video shallow features respectively output by the plurality of moving window-based multi-head attention feature extraction submodules.
[0058] In an optional example, the multiple moving window-based multi-head attention feature extraction submodules are composed of L moving window-based Swin transformer blocks, and each Swin transformer block includes: window multi-head attention (W-MSA), layer normalization (LN) and residual connection. The output of the lth layer (i.e., the lth Swin transformer block) is: .in, Include:
[0059] ;
[0060] Among them, MLP is a multi-layer perceptron, is the output of the l-1th layer (i.e. the l-1th Swin transformer block), is the intermediate output feature of the lth layer (i.e., the lth Swin transformer block).
[0061] This continues until the output of the Lth layer is obtained as the final shallow feature extracted . m represents the mode, Represent satellite, drone, and ground tower modes respectively.
[0062] After complete shallow layer extraction, the shallow features of satellite video, drone video, and ground tower video of the three modalities are:
[0063] .
[0064] Step S25: splicing the shallow features of the satellite video, the shallow features of the drone video, and the shallow features of the ground tower video in the channel dimension to obtain splicing features, and using a multi-layer perceptron to process the splicing features to obtain the multimodal fusion shallow features.
[0065] In this embodiment, the shallow features of satellite video, drone video, and ground tower video can be spliced in the channel dimension to obtain spliced features. The spliced features are then processed using a multi-layer perceptron (MLP) (for example, dimensionality reduction and feature fusion) to obtain multimodal fusion shallow features. The multimodal fusion shallow features are expressed as .
[0066] In combination with any of the above embodiments, the present invention also provides a multimodal video fusion method for wide area visual network, referring to Figure 2 , Figure 2 : is a flowchart of a multimodal video fusion method for wide-area visual networking according to another embodiment of the present invention. In this method, the above step S14 may specifically include the following steps S31 to S32:
[0067] Step S31: aligning different groups using the optical flow information between the different groups through the fusion feature circular refinement module to obtain aligned features.
[0068] In this embodiment, the fusion feature loop refinement module includes multiple layers, each with the same structure. A global loop of feature propagation is performed on each layer: feature refinement is performed on each group of multimodal fusion shallow features in chronological order on each layer, so that information interaction occurs between the video frame features of different groups of multimodal fusion shallow features. When performing global loop feature refinement through the fusion feature loop refinement module, this embodiment first uses the optical flow information between different groups of multimodal fusion shallow features to align the different groups of multimodal fusion shallow features to obtain aligned features.
[0069] Step S32: Processing the same group of multimodal refined features at different layers, the multiple groups of multimodal fused shallow features, and the aligned features through the fusion feature loop refinement module to obtain the multiple groups of multimodal refined features.
[0070] In this embodiment, multiple sets of multimodal fusion shallow features are sequentially subjected to multi-layer global cyclic feature refinement processing to obtain multimodal refined features at different layers corresponding to the multiple sets of multimodal fusion shallow features (i.e., multiple sets of multimodal refined features at different layers). After obtaining the aligned features, the fused feature cyclic refinement module can be used to process the same set of multimodal refined features at different layers, the multiple sets of multimodal fusion shallow features, and the aligned features to obtain multiple sets of multimodal refined features output by the fused feature cyclic refinement module.
[0071] With respect to local parallel feature refinement, each group of multimodal refined features at each layer in this embodiment is obtained by performing parallel feature refinement within the group.
[0072] In combination with any of the above embodiments, in one embodiment, the present invention further provides a multimodal video fusion method for wide-area visual networking. In this method, the above-mentioned step S14 of "performing parallel feature refinement within each group by the fusion feature cyclic refinement module, and performing global cyclic feature refinement between different groups by the fusion feature cyclic refinement module to obtain multiple groups of multimodal refined features" can specifically include the following steps S41 to S42:
[0073] Step S41: aligning the t-1th group with the tth group using the optical flow information between the t-1th group and the tth group through the fusion feature circular refinement module to obtain aligned features of the t-1th group aligned with the tth group.
[0074] In this embodiment, the number of groups of multimodal fusion shallow features is T, where T is an integer greater than 2. The fusion feature cyclic refinement module includes multiple layers, each with the same structure. A global cyclic feature propagation is performed on each layer: each group of multimodal fusion shallow features is refined in chronological order on each layer, thereby enabling information exchange between video frame features between different groups of multimodal fusion shallow features.
[0075] In this embodiment, when performing global cyclic feature refinement through the fusion feature cyclic refinement module, the optical flow information between the t-1th group and the tth group is used to align the t-1th group with the tth group. Specifically, the t-1th group is aligned with the tth group to obtain aligned features in which the t-1th group is aligned with the tth group. Where t is an integer in the range [2, T].
[0076] Step S42: The tth group of multimodal refined features at different layers, the tth group of multimodal fusion shallow features, and the aligned features of the t-1th group aligned to the tth group are processed through the fusion feature loop refinement module to obtain the tth group of multimodal refined features.
[0077] In this embodiment, for the tth group of multimodal fusion shallow features among the multiple groups of multimodal fusion shallow features, after undergoing multi-layer global cyclic feature refinement processing in sequence, the multimodal refinement features of different layers corresponding to the tth group of multimodal fusion shallow features (i.e., the multimodal refinement features of the tth group at different layers) can be obtained. After obtaining the aligned features aligned with the tth group by the t-1th group, the tth group of multimodal refinement features at different layers, the tth group of multimodal fusion shallow features, and the aligned features aligned with the tth group by the fusion feature cyclic refinement module can be processed to obtain the tth group of multimodal refinement features, and then multiple groups of multimodal refinement features are output.
[0078] With respect to local parallel feature refinement, the t-th group of multimodal refined features at each layer in this embodiment is obtained by performing parallel feature refinement within the t-th group.
[0079] In conjunction with any of the above embodiments, in one embodiment, the present invention further provides a multimodal video fusion method for wide-area visual networking. In this method, the fusion feature loop refinement module includes I layers, where I is an integer greater than 2. Furthermore, the above S41 may specifically include step S51:
[0080] Step S51: Refine the multimodal features of the t-1th group at the i-1th layer As the key vector, the multimodal refinement feature of the tth group at the i-1 layer As the query vector, take the multimodal refinement features of the t-1th group at the i-th layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to perform guided deformable attention processing using the optical flow information between the t-1 group and the t group to obtain the aligned features of the t-1 group aligned to the t group at the i layer. .
[0081] In this embodiment, for the multimodal fusion shallow features of the tth (t∈{2,…,N})th group of the ith (i∈{2,…,I})th layer, the refinement of its features requires the use of the multimodal refinement features of the t-1th group at the i-1th layer. , the multimodal refined features of the tth group at the i-1 layer And the multimodal refinement features of the t-1th group at layer i .
[0082] Specifically, it can be the multimodal refinement feature of the t-1th group at the i-1th layer As the key vector, the multimodal refinement feature of the tth group at the i-1 layer As the query vector, take the multimodal refinement features of the t-1th group at the i-th layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to utilize the optical flow information between the t-1th group and the tth group. , guide the deformable attention process to obtain the aligned features of the t-1th group to the tth group at the i-th layer .
[0083] Among them, for local parallel feature refinement, the tth group of multimodal refinement features at the i-1 layer It is obtained by parallel feature refinement in group t; the multimodal refinement feature of group t-1 at layer i-1 is obtained by parallel feature refinement in group t-1 at layer i-1. and the multimodal refined features of the t-1th group at layer i They are obtained by performing parallel feature refinement on different layers within the t-1th group.
[0084] In an optional example, the aligned features of the t-1th group aligned to the tth group at the i-th layer can be obtained by the following formula: :
[0085] ;
[0086] Among them, the As the value vector V, the semicolon after and As key vector K and query vector Q respectively, based on Q, K and V Guided deformable attention is performed. GDA is a special attention mechanism that first performs pre-alignment based on optical flow, then predicts the offset of the optical flow and samples related features based on the optical flow and the offset of the optical flow, calculates the attention weight, and finally realizes the interaction of video frame features between different groups.
[0087] In an optional example, in the first layer, the aligned features of the t-1th group aligned to the tth group in the first layer are obtained. The method is as follows:
[0088] Specifically, it can be the multimodal fusion shallow features of the t-1th group at the 0th layer (i.e., the t-1th group of multimodal fusion shallow features) as the key vector, and the tth group of multimodal fusion shallow features at layer 0 as the query vector (i.e., the tth group of multimodal fusion shallow features), with the t-1th group of multimodal refinement features in the first layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to utilize the optical flow information between the t-1th group and the tth group. , and the guided deformable attention processing is performed to obtain the aligned feature of the t-1th group to the tth group in the 1st layer .
[0089] In combination with any of the above embodiments, in an implementation, the application further provides a multi-modal video fusion method for a wide-area visual internet. In the method, the step S42 can specifically include the following step S61:
[0090] Step S61: The RFR module in the fusion feature cycle refinement module is used to process the multi-modal refined features of the tth group in the 1st layer to the i-1th layer, the multi-modal fusion shallow layer features of the tth group, and the aligned feature of the t-1th group to the tth group in the i-th layer to obtain the multi-modal refined features of the tth group in the i-th layer.
[0091] In the embodiment, the aligned feature of the t-1th group to the tth group in the i-th layer and the multi-modal refined features of the tth group in different layers are comprehensively used to generate the multi-modal refined features of the tth group in the i-th layer. Specifically, the RFR module (i.e., the cycle feature refinement module) in the fusion feature cycle refinement module is used to process the multi-modal refined features of the tth group in the 1st layer to the i-1th layer, the multi-modal fusion shallow layer features of the tth group, and the aligned feature of the t-1th group to the tth group in the i-th layer to obtain the multi-modal refined features of the tth group in the i-th layer, until the multi-modal refined features of the tth group in the I-th layer (i.e., the multi-modal refined features of the tth group output by the fusion feature cycle refinement module) are obtained, and then the multi-modal refined features of each group in the I-th layer (i.e., the multi-modal refined features of multiple groups output by the fusion feature cycle refinement module) are obtained.
[0092] In an optional example, the multi-modal refined features of the tth group in the i-th layer can be obtained by the following formula:
[0093] ;
[0094] wherein, is the multi-modal fusion shallow layer features of the tth group, is the multi-modal refined features of the tth group in the 1st layer to the i-1th layer, is the aligned feature of the t-1th group to the tth group in the i-th layer, and RFR is the cycle feature refinement module used to fuse different layer features and the aligned features.
[0095] In an optional example, in the first layer, the multi-modal refined features of the tth group in the 1st layer can be obtained by the following method:
[0096] After obtaining the aligned features of the t-1th group aligned to the tth group at the first layer After that, the t-th group of multimodal fusion shallow features can be and the aligned features of the t-1th group aligned to the tth group at the first layer Input the RFR module in the fusion feature loop refinement module for processing to obtain the t-th group of multimodal refinement features in the first layer For example, the multimodal refinement feature of the tth group at the first layer can be obtained by the following formula: :
[0097] =RFR( , ).
[0098] In an optional example, for the first set of multimodal fusion shallow features, how to calculate the corresponding refined features at each layer:
[0099] In each layer, the multimodal fusion shallow features of group 1 are And the first group of multimodal refined features of all layers before this layer are input into the RFR module in the fusion feature loop refinement module to obtain the first group of multimodal refined features in this layer.
[0100] For example: for the first layer, the multimodal fusion shallow features of the first group Input the RFR module in the fusion feature loop refinement module to obtain the first set of multimodal refinement features at the first layer :
[0101] =RFR( ).
[0102] For the second layer, the multimodal fusion shallow features of the first group are and the multimodal refinement features of group 1 at layer 1 Input the RFR module in the fusion feature loop refinement module to obtain the first group of multimodal refinement features in the second layer :
[0103] =RFR( , ).
[0104] For the i-th layer (i∈[3, I]), the multimodal fusion shallow features of group 1 are , the first group of multimodal refinement features at the first layer , ..., the first group of multimodal refinement features at layer i-1 Input the RFR module in the fusion feature loop refinement module to obtain the first group of multimodal refinement features at the i-th layer :
[0105] =RFR( , ,…, ).
[0106] It should be noted that the first group here is alternating because this embodiment needs to alternate the order of the video sequence. When the detailed features of different groups of videos are generated in a forward order, the video group at the front end of the video is the first group. When the detailed features of different groups of videos are generated in a reverse order, the video group at the end of the video is the first group.
[0107] In conjunction with any of the above embodiments, in one implementation, the present invention further provides a multimodal video fusion method for wide-area visual networking. In this method, before step S15, step S71 may be further included, and the "upsampling the multiple sets of multimodal refined features to obtain a reconstructed video" in step S15 may specifically include step S72:
[0108] Step S71: input the multiple groups of multimodal refined features into multiple moving window-based multi-head attention feature extraction sub-modules in the feature refinement module for further feature refinement to obtain multiple groups of multimodal refined features.
[0109] In this embodiment, the multimodal video fusion network also includes a feature refinement module for further refining the multiple sets of multimodal features that have undergone global loop and local parallel feature refinement processing. The feature refinement module includes multiple moving window-based multi-head attention feature extraction submodules. The multiple sets of multimodal features can be input into the multiple moving window-based multi-head attention feature extraction submodules in the feature refinement module for further feature refinement processing to obtain multiple sets of multimodal refined features.
[0110] Among them, the multiple moving window-based multi-head attention feature extraction sub-modules in the feature refinement module are the same as or similar to the multiple moving window-based multi-head attention feature extraction sub-modules in the multimodal shallow feature extraction module, and the processing process of the multiple moving window-based multi-head attention feature extraction sub-modules can refer to the aforementioned embodiments.
[0111] Step S72: up-sampling the multiple groups of multimodal refined features to obtain the reconstructed video.
[0112] In this embodiment, after obtaining multiple sets of multimodal refined features, the multiple sets of multimodal refined features can be input into a video reconstruction module, which then upsamples the multiple sets of multimodal refined features to produce a reconstructed video. In an optional example, the video reconstruction module includes at least a pixel shuffling upsampling module, which can upsample the multiple sets of multimodal refined features to produce a reconstructed video.
[0113] In addition, in an optional embodiment, in order to implement the grouping operation in the above step S13, the final shallow features (i.e., satellite video shallow features, drone video shallow features, or ground tower video shallow features) can be After deformation, the characteristics are , is the number of features within a group, thereby achieving grouping. Based on this, before step S71, multiple groups of multimodal refinement features can be transformed into their original shape , followed by feature refinement and video reconstruction.
[0114] In conjunction with any of the above embodiments, in one implementation, after obtaining multimodal video data, before inputting the multimodal video data into the multimodal shallow feature extraction module, the video data of different modalities needs to be spatially segmented and aligned. Specifically, because the resolution and field of view of video data of different modalities are different, it is necessary to segment the satellite video data with lower resolution and larger field of view so that the area covered by the satellite video data roughly overlaps with the drone video data and the ground tower video data (e.g., the feature similarity of the covered area is greater than a first threshold, and / or the area overlap of the covered area is greater than a second threshold). The multimodal video data consisting of the segmented satellite video data, drone video data, and ground tower video data is then input into the multimodal shallow feature extraction module.
[0115] In combination with any of the above embodiments, in one implementation, after obtaining multimodal video data, before inputting the multimodal video data into the multimodal shallow feature extraction module, the perspective differences caused by different shooting angles in the satellite video data, drone video data, and ground tower video data are processed (such as adjusting the perspective range, etc.), and the satellite video data, drone video data, and ground tower video data after eliminating the perspective differences are input into the multimodal shallow feature extraction module.
[0116] In combination with any of the above embodiments, in one implementation, the multimodal video fusion network pre-trained in this embodiment is obtained by training an initial multimodal video fusion network, which includes at least: an initial multimodal shallow feature extraction module, an initial fusion feature loop refinement module, and an initial video reconstruction module, wherein the first feature extraction submodule, the second feature extraction submodule, and the third feature extraction submodule in the initial multimodal shallow feature extraction module are feature extraction submodules corresponding to satellite video data, drone video data, and ground tower video data, respectively. The initial multimodal shallow feature extraction module also includes multiple initial multi-head attention feature extraction submodules based on moving windows, and the initial fusion feature loop refinement module includes an initial guided deformable attention submodule and an initial RFR module.
[0117] In this embodiment, based on sample multimodal video data including sample satellite video data, sample drone video data and sample ground tower video data, the sample multimodal video data can be input into the initial multimodal video fusion network, and after being processed by the initial multimodal shallow feature extraction module, the initial fusion feature loop refinement module and the initial video reconstruction module respectively, the sample reconstructed video output by the initial multimodal video fusion network is obtained; then, based on the semantic features corresponding to the sample reconstructed video and the semantic features corresponding to the sample multimodal video data, the loss is calculated, and based on the loss value, the network parameters of the first feature extraction submodule, the second feature extraction submodule, the third feature extraction submodule, multiple initial multi-head attention feature extraction submodules based on moving windows, the initial guided deformable attention submodule, the initial RFR module and the initial video reconstruction module in the initial multimodal shallow feature extraction module are updated until the loss converges to obtain a trained multimodal video fusion network.
[0118] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required for the embodiments of the present invention.
[0119] Based on the same inventive concept, an embodiment of the present invention provides a multimodal video fusion device for wide area visual networking. Figure 3 , Figure 3 This is a structural block diagram of a multimodal video fusion device for wide area visual networking provided by one embodiment of the present invention. Figure 3 As shown, the device includes:
[0120] The data acquisition module is configured to obtain multi-modal video data by synchronously shooting a target scene through a camera of a remote sensing satellite, a camera of an unmanned aerial vehicle, and a camera of a ground tower, wherein the multi-modal video data comprises satellite video data, unmanned aerial vehicle video data, and ground tower video data.
[0121] The feature extraction module is configured to input the multi-modal video data into a multi-modal shallow feature extraction module to obtain multi-modal fusion shallow features.
[0122] The feature grouping module is configured to group process the multi-modal fusion shallow features in a time sequence to obtain a plurality of groups of multi-modal fusion shallow features, wherein each group of multi-modal fusion shallow features comprises N multi-modal fusion shallow features that are continuous in time, and N is an integer greater than 1.
[0123] The feature refinement module is configured to input the plurality of groups of multi-modal fusion shallow features into a fusion feature cyclic refinement module, perform parallel feature refinement in each group through the fusion feature cyclic refinement module, and perform global cyclic feature refinement between different groups through the fusion feature cyclic refinement module to obtain a plurality of groups of multi-modal refined features.
[0124] The video reconstruction module is configured to input the plurality of groups of multi-modal refined features into a video reconstruction module, perform up-sampling on the plurality of groups of multi-modal refined features, and obtain a reconstructed video, wherein the reconstructed video fuses information provided by the satellite video data, the unmanned aerial vehicle video data, and the ground tower video data.
[0125] Optionally, the feature extraction module comprises:
[0126] The first extraction module is configured to input the satellite video data into a first feature extraction submodule in the multi-modal shallow feature extraction module to obtain a convolutional satellite video feature.
[0127] The second extraction module is configured to input the unmanned aerial vehicle video data into a second feature extraction submodule in the multi-modal shallow feature extraction module to obtain a convolutional unmanned aerial vehicle video feature.
[0128] The third extraction module is configured to input the ground tower video data into a third feature extraction submodule in the multi-modal shallow feature extraction module to obtain a convolutional ground tower video feature.
[0129] The fourth extraction module is configured to input the convolutional satellite video feature, the convolutional unmanned aerial vehicle video feature, and the convolutional ground tower video feature into a plurality of multi-head attention feature extraction submodules based on a moving window, respectively, to obtain a satellite video shallow feature, an unmanned aerial vehicle video shallow feature, and a ground tower video shallow feature.
[0130] a splicing module for splicing the shallow features of the satellite video, the shallow features of the drone video, and the shallow features of the ground tower video in the channel dimension to obtain splicing features, and processing the splicing features using a multi-layer perceptron to obtain the multimodal fusion shallow features;
[0131] Among them, the first parameter value of the convolution kernel in the first feature extraction submodule is adapted to the video resolution of the satellite video data, the second parameter value of the convolution kernel in the second feature extraction submodule is adapted to the video resolution of the drone video data, and the third parameter value of the convolution kernel in the third feature extraction submodule is adapted to the video resolution of the ground tower video data.
[0132] Optionally, the fusion feature loop refinement module includes multiple layers; the feature refinement module includes:
[0133] A first alignment module is configured to perform alignment between different groups by looping through the fusion feature refinement module and utilizing optical flow information between different groups to obtain aligned features;
[0134] The first refinement module is used to process the same group of multimodal refinement features at different layers, the multiple groups of multimodal fusion shallow features and the aligned features through the fusion feature loop refinement module to obtain the multiple groups of multimodal refinement features; wherein each group of multimodal refinement features at each layer is obtained by parallel feature refinement within the group.
[0135] Optionally, the fusion feature loop refinement module includes multiple layers, and the number of the multiple groups of multimodal fusion shallow features is T; the feature refinement module includes:
[0136] A second alignment module is configured to align the t-1th group with the tth group by using the optical flow information between the t-1th group and the tth group through the fusion feature circular refinement module to obtain an aligned feature in which the t-1th group is aligned with the tth group;
[0137] A second refinement module is configured to process the tth group of multimodal refinement features at different layers, the tth group of multimodal fusion shallow features, and the aligned features of the t-1th group aligned to the tth group through the fusion feature circular refinement module to obtain the tth group of multimodal refinement features;
[0138] The multimodal refined features of the t-th group at each layer are obtained by performing parallel feature refinement within the t-th group, and T is an integer greater than 2.
[0139] Optionally, the fusion feature loop refinement module includes I layer, a second alignment module, including:
[0140] The second alignment submodule is used to refine the multimodal features at the i-1 layer with the t-1th group As the key vector, the multimodal refinement feature of the tth group at the i-1 layer As the query vector, take the multimodal refinement features of the t-1th group at the i-th layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to perform guided deformable attention processing using the optical flow information between the t-1 group and the t group to obtain the aligned features of the t-1 group aligned to the t group at the i layer. ;
[0141] Among them, the multimodal refinement features of the tth group at the i-1 layer It is obtained by parallel feature refinement in group t; the multimodal refinement feature of group t-1 at layer i-1 is obtained by parallel feature refinement in group t-1 at layer i-1. and the multimodal refined features of the t-1th group at layer i are obtained by performing parallel feature refinement on different layers within the t-1th group; I is an integer greater than 2.
[0142] Optionally, the second refinement module includes:
[0143] The second refinement submodule is used to refine the multimodal refinement features of the t-th group in the 1st to i-1th layers, the t-th group of multimodal fusion shallow features, and the aligned features of the t-1th group aligned to the t-th group in the i-th layer through the RFR module in the fusion feature cycle refinement module. Processing is performed to obtain the t-th group of multimodal refined features at the i-th layer.
[0144] Optionally, the device further comprises:
[0145] A feature refinement module is configured to input the multiple sets of multimodal refined features into multiple moving window-based multi-head attention feature extraction submodules in the feature refinement module before inputting the multiple sets of multimodal refined features into the video reconstruction module for further feature refinement to obtain multiple sets of multimodal refined features;
[0146] Video reconstruction module, including:
[0147] The video reconstruction submodule is used to upsample the multiple groups of multimodal refined features to obtain the reconstructed video.
[0148] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multimodal video fusion method for wide-area visual networking as described in any of the above embodiments of the present invention are implemented.
[0149] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as Figure 4 As shown, Figure 4 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the steps of the multimodal video fusion method for wide-area visual networking described in any of the above embodiments of the present invention.
[0150] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0151] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0152] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0153] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0154] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0156] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0157] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0158] The above is a detailed introduction to the multimodal video fusion method, device, equipment and medium for wide-area visual networking provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A multimodal video fusion method for wide area visual networking, characterized in that: The method comprises: The target scene is captured synchronously by cameras of remote sensing satellites, drones, and ground towers to obtain multimodal video data, wherein the multimodal video data includes: satellite video data, drone video data, and ground tower video data; Inputting the multimodal video data into a multimodal shallow feature extraction module to obtain multimodal fusion shallow features; The multimodal fusion shallow features are grouped and processed in chronological order to obtain multiple groups of multimodal fusion shallow features, each group of multimodal fusion shallow features includes N multimodal fusion shallow features that are continuous in time, where N is an integer greater than 1; Inputting the multiple groups of multimodal fused shallow features into a fusion feature cyclic refinement module, performing parallel feature refinement within each group through the fusion feature cyclic refinement module, and performing global cyclic feature refinement between different groups through the fusion feature cyclic refinement module, to obtain multiple groups of multimodal refined features; Inputting the multiple sets of multimodal refined features into a video reconstruction module, upsampling the multiple sets of multimodal refined features to obtain a reconstructed video, wherein the reconstructed video fuses information provided by the satellite video data, the drone video data, and the ground tower video data; The fusion feature cyclic refinement module includes multiple layers, and the number of the multiple groups of multimodal fusion shallow features is T; the fusion feature cyclic refinement module performs parallel feature refinement in each group, and the fusion feature cyclic refinement module performs global cyclic feature refinement between different groups to obtain multiple groups of multimodal refined features, including: By using the fusion feature circular refinement module, the optical flow information between the t-1th group and the tth group is used to align the t-1th group and the tth group, thereby obtaining an aligned feature in which the t-1th group is aligned with the tth group; The tth group of multimodal refined features at different layers, the tth group of multimodal fused shallow features, and the aligned features of the t-1th group aligned to the tth group are processed by the fusion feature loop refinement module to obtain the tth group of multimodal refined features; The multimodal refined features of the tth group at each layer are obtained by performing parallel feature refinement within the tth group, and T is an integer greater than 2; The fused feature cyclic refinement module includes a layer I, through which the optical flow information between the t-1th group and the tth group is used to align the t-1th group and the tth group, to obtain aligned features of the t-1th group aligned to the tth group, including: Refine the multimodal features of the t-1th group at the i-1th layer As the key vector, the multimodal refinement feature of the tth group at the i-1 layer As the query vector, take the multimodal refinement features of the t-1th group at the i-th layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to perform guided deformable attention processing using the optical flow information between the t-1 group and the t group to obtain the aligned features of the t-1 group aligned to the t group at the i layer. ; Among them, the multimodal refinement features of the tth group at the i-1 layer It is obtained by parallel feature refinement in group t; the multimodal refinement feature of group t-1 at layer i-1 is obtained by parallel feature refinement in group t-1 at layer i-1. and the multimodal refined features of the t-1th group at layer i are obtained by performing parallel feature refinement on different layers within the t-1th group; I is an integer greater than 2.
2. The multimodal video fusion method for wide area visual networking according to claim 1, characterized in that: Inputting the multimodal video data into a multimodal shallow feature extraction module to obtain multimodal fusion shallow features, including: Inputting the satellite video data into the first feature extraction submodule in the multimodal shallow feature extraction module to obtain convolved satellite video features; Inputting the drone video data into the second feature extraction submodule in the multimodal shallow feature extraction module to obtain convolved drone video features; Inputting the ground tower video data into the third feature extraction submodule in the multimodal shallow feature extraction module to obtain convolved ground tower video features; The convolved satellite video features, the convolved drone video features, and the convolved ground tower video features are respectively input into a plurality of moving window-based multi-head attention feature extraction submodules to obtain satellite video shallow features, drone video shallow features, and ground tower video shallow features; The shallow features of the satellite video, the shallow features of the drone video, and the shallow features of the ground tower video are spliced in the channel dimension to obtain spliced features, and the spliced features are processed using a multi-layer perceptron to obtain the multimodal fusion shallow features; Among them, the first parameter value of the convolution kernel in the first feature extraction submodule is adapted to the video resolution of the satellite video data, the second parameter value of the convolution kernel in the second feature extraction submodule is adapted to the video resolution of the drone video data, and the third parameter value of the convolution kernel in the third feature extraction submodule is adapted to the video resolution of the ground tower video data.
3. The multimodal video fusion method for wide area visual networking according to claim 1, characterized in that: The fusion feature cyclic refinement module includes multiple layers; the multiple groups of multimodal fusion shallow features are input into the fusion feature cyclic refinement module, the fusion feature cyclic refinement module performs parallel feature refinement in each group, and the fusion feature cyclic refinement module performs global cyclic feature refinement between different groups to obtain multiple groups of multimodal refined features, including: By using the fusion feature cycle refinement module, the optical flow information between different groups is used to align the different groups to obtain aligned features; The fusion feature loop refinement module processes the same group of multimodal refinement features at different layers, the multiple groups of multimodal fusion shallow features, and the aligned features to obtain the multiple groups of multimodal refinement features; wherein, each group of multimodal refinement features at each layer is obtained by performing parallel feature refinement within the group.
4. The multimodal video fusion method for wide area visual networking according to claim 1, characterized in that: The tth group of multimodal refined features at different layers, the tth group of multimodal fused shallow features, and the aligned features of the t-1th group aligned to the tth group are processed by the fusion feature cyclic refinement module to obtain the tth group of multimodal refined features, including: Through the RFR module in the fusion feature cycle refinement module, the multimodal refined features of the tth group in the 1st layer to the i-1th layer, the tth group of multimodal fusion shallow features and the aligned features of the t-1th group to the tth group in the i-th layer are aligned. Processing is performed to obtain the t-th group of multimodal refined features at the i-th layer.
5. The multimodal video fusion method for wide area visual networking according to any one of claims 1 to 4, characterized in that: Before inputting the multiple sets of multimodal refined features into the video reconstruction module, the method further includes: Inputting the multiple sets of multimodal refined features into multiple moving window-based multi-head attention feature extraction submodules in the feature refinement module for further feature refinement to obtain multiple sets of multimodal refined features; Upsampling the multiple sets of multimodal refined features to obtain a reconstructed video includes: The multiple groups of multimodal refined features are up-sampled to obtain the reconstructed video.
6. A multimodal video fusion device for wide-area visual networking, characterized in that: The device comprises: A data acquisition module is used to synchronously capture the target scene through cameras of remote sensing satellites, drones, and ground towers to obtain multimodal video data, wherein the multimodal video data includes: satellite video data, drone video data, and ground tower video data; A feature extraction module, configured to input the multimodal video data into a multimodal shallow feature extraction module to obtain multimodal fusion shallow features; A feature grouping module is used to group the multimodal fusion shallow features in chronological order to obtain multiple groups of multimodal fusion shallow features, each group of multimodal fusion shallow features includes N multimodal fusion shallow features that are continuous in time, where N is an integer greater than 1; A feature refinement module is configured to input the plurality of groups of multimodal fused shallow features into a fusion feature cyclic refinement module, perform parallel feature refinement within each group through the fusion feature cyclic refinement module, and perform global cyclic feature refinement between different groups through the fusion feature cyclic refinement module to obtain a plurality of groups of multimodal refined features; a video reconstruction module, configured to input the multiple sets of multimodal refined features into the video reconstruction module, upsample the multiple sets of multimodal refined features, and obtain a reconstructed video, wherein the reconstructed video incorporates information provided by the satellite video data, the drone video data, and the ground tower video data; The fusion feature loop refinement module includes multiple layers, and the number of the multiple groups of multimodal fusion shallow features is T; the feature refinement module includes: A second alignment module is configured to align the t-1th group with the tth group by using the optical flow information between the t-1th group and the tth group through the fusion feature circular refinement module to obtain an aligned feature in which the t-1th group is aligned with the tth group; A second refinement module is configured to process the tth group of multimodal refinement features at different layers, the tth group of multimodal fusion shallow features, and the aligned features of the t-1th group aligned to the tth group through the fusion feature circular refinement module to obtain the tth group of multimodal refinement features; The multimodal refined features of the tth group at each layer are obtained by performing parallel feature refinement within the tth group, and T is an integer greater than 2; The fusion feature loop refinement module includes I layer, the second alignment module includes: The second alignment submodule is used to refine the multimodal features at the i-1 layer with the t-1th group As the key vector, the multimodal refinement feature of the tth group at the i-1 layer As the query vector, take the multimodal refinement features of the t-1th group at the i-th layer As a value vector, the guided deformable attention submodule in the fusion feature loop refinement module is used to perform guided deformable attention processing using the optical flow information between the t-1 group and the t group to obtain the aligned features of the t-1 group aligned to the t group at the i layer. ; Among them, the multimodal refinement features of the tth group at the i-1 layer It is obtained by parallel feature refinement in group t; the multimodal refinement feature of group t-1 at layer i-1 is obtained by parallel feature refinement in group t-1 at layer i-1. and the multimodal refined features of the t-1th group at layer i are obtained by performing parallel feature refinement on different layers within the t-1th group; I is an integer greater than 2.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the multimodal video fusion method for wide-area visual networking is implemented as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal video fusion method for wide-area visual networking is implemented as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Video super-division model training method and device and video super-division processing method and device
CN115631093A
Multi-modal image fusion method based on feature information interaction
CN116071281A