Multi-frame adaptive fusion panoramic video feature alignment method

Through the multi-frame adaptive fusion method, the feature alignment of panoramic video frames is used to use splicing and convolution technology to solve the problems of uneven distribution and flexible differences in feature alignment of panoramic video frames, and efficient feature alignment and multi-frame feature fusion are achieved.

CN120088147APending Publication Date: 2025-06-03CHINA STATE SHIPBUILDING CORP NO 707 RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510128509.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Panoramic video frame feature alignment faces the problems of uneven distribution of panoramic video images and flexible and changeable differences in multiple frame images, and the prior art is difficult to achieve efficient feature alignment.

Method used

Adaptive fusion method of multi-frames is adopted to splice the features of reference frames and input frames, and use standard convolutions, activation functions and hollow convolutions to obtain the offset matrix, and combine deformable convolutions to perform recursive alignment to realize the alignment of the features of the panoramic video input frames to the reference frame.

Benefits of technology

It effectively expands the receptive field of feature data, realizes the effective fusion of multi-frame image features, solves the problems of uneven distribution and flexible differences in feature alignment of panoramic videos, and realizes efficient panoramic video feature alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088147A_ABST
    Figure CN120088147A_ABST
Patent Text Reader

Abstract

The invention relates to a panoramic video feature alignment method based on multi-frame adaptive fusion. The panoramic video feature alignment method comprises the following steps: step 1), splicing features of a reference frame and an input frame; step 2), performing standard convolution, activation function and cavity convolution on the spliced features to obtain an offset matrix of the input frame features relative to the reference frame features in space; step 3), according to the offset provided by the offset matrix, carrying out deformable convolution on the reference frame feature map to obtain a feature map after preliminary alignment; and step 4), recursively performing the step 1), the step 2) and the step 3) for several times to obtain a reference frame feature map after final alignment. According to the panoramic video feature alignment method based on multi-frame adaptive fusion, a larger and more flexible receptive field is provided for feature data in feature alignment, and meanwhile, the features of multiple frames of images are effectively fused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of feature alignment between video frames, and particularly relates to a panoramic video feature alignment method for multi-frame adaptive fusion. Background Art

[0002] Feature alignment of images and videos has wide applications in fields such as multi-frame fusion of images, video anti-shake, image stitching, and super-resolution. The rapid popularization of panoramic video stitching and super-resolution technologies has given rise to the need for feature alignment of panoramic video frames. How to achieve efficient feature alignment of panoramic videos has become an important research issue. The following is an introduction to the relevant background art.

[0003] A panoramic image maps the scene within a 360° range onto a plane through a geometric mapping relationship, and then through the correction process of a panoramic player, a three-dimensional visual effect is achieved. Such an image has the characteristics of all-round, real-scene, and 360-degree panoramic view, enabling the observer to feel the surrounding environment as if on the spot. A panoramic video is a video shot omnidirectionally at 360 degrees with a 3D camera, and users can freely adjust the video up, down, left, and right for viewing. The panoramic video breaks the perspective fixity of traditional 2D videos and gives the viewing option entirely to the audience, enabling the audience to explore every angle in the video more freely. Panoramic images and panoramic videos have wide applications in fields such as tourism exhibitions, urban introductions, medical observations, and entertainment.

[0004] Feature alignment of video frames refers to ensuring the accurate matching and correspondence of image features between consecutive video frames through a series of algorithms and technical means. This alignment process is crucial for maintaining the coherence, stability, and high-quality visual effect of the video. When performing feature alignment, it is first necessary to extract key feature information from each frame, such as corner points, edges, textures, etc. These feature information can be automatically detected and extracted through computer vision algorithms. Next, by comparing the feature points between adjacent frames, their relative positions and transformation relationships can be determined. The goal of feature alignment is to find a transformation method that enables the feature points between adjacent frames to correspond one by one. This usually involves image registration technology, that is, calculating a transformation matrix to align one frame of image with another frame of image, and the transformation matrix describes how to map the pixel positions in one frame to the corresponding positions in another frame.

[0005] Each frame of the panoramic video is a panoramic image. Since the panoramic image displays the information of a spherical space in a plane, the real scene is not evenly distributed in the panoramic video frame, which brings difficulties to the feature alignment of the panoramic video. To effectively address the uneven distribution of the panoramic video frame and the flexible differences between multiple frames of images, a feature alignment method for panoramic video with multi-frame adaptive fusion is proposed to align the features of the input frame of the panoramic video to the reference frame. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the existing panoramic video feature alignment technology, and provide a feature alignment method for panoramic video with multi-frame adaptive fusion, which provides a larger and more flexible receptive field for the feature data in feature alignment, and effectively fuses the features of multiple frames of images at the same time.

[0007] The present invention solves its technical problems through the following technical solutions:

[0008] A feature alignment method for panoramic video with multi-frame adaptive fusion, characterized in that it includes the following steps:

[0009] Step 1): Concatenate the features of the reference frame and the input frame;

[0010] Step 2): The concatenated features pass through a standard convolution, an activation function, and a dilated convolution to obtain an offset matrix of the input frame features relative to the reference frame features in space;

[0011] Step 3): According to the offset provided by the offset matrix, the reference frame feature map passes through a deformable convolution to obtain a preliminarily aligned feature map;

[0012] Step 4): Recursively perform Step 1), Step 2), and Step 3) several times to obtain the finally aligned reference frame feature map.

[0013] Moreover, the specific process of concatenating the features of the reference frame and the input frame in Step 1) is as follows: Obtain the reference frame feature map F ref and the feature map F nbor of its input frame. It is required that the sizes of the reference frame and the input frame feature maps are exactly the same, and the feature maps of the reference frame and the input frame are concatenated along the channel direction into a new feature map.

[0014] Moreover, the specific process of obtaining the offset matrix of the input frame features relative to the reference frame features in space by passing the concatenated features through a standard convolution, an activation function, and a dilated convolution in Step 2) is as follows:

[0015] Step 2.1 The concatenated feature map passes through a standard convolution with a convolution kernel size of 3×3;

[0016] Step 2.2 The feature map passes through an activation function, and the activation function is the Relu activation function. The expression of the activation function is as follows:

[0017]

[0018] Step 2.3 The feature map passes through dilated convolution. Let the convolution kernel matrix be K m×n , and the element at the (i, j) position in the output offset matrix O is calculated as follows:

[0019]

[0020] where a is the dilation rate (hole rate). The larger a is, the larger the receptive field of the original video frame obtained by the output feature is.

[0021] Moreover, the specific process of obtaining the preliminarily aligned feature map by performing deformable convolution on the reference frame feature map according to the offset provided by the offset matrix in step 3) is as follows:

[0022] According to the offset provided by the offset matrix O, the reference frame feature map F ref passes through deformable convolution to obtain the preliminarily aligned feature map F align , and let the convolution kernel matrix of the deformable convolution be A m×n , and the element at the (i, j) position in F align is calculated as follows:

[0023] F align (i, j) = ∑ m ∑ n F ref [i + O(i, j) - m, j + O(i, j) - n]A(m, n).

[0024] Moreover, the specific process of recursively performing step 1), step 2), and step 3) several times in step 4) to obtain the finally aligned reference frame feature map is as follows: Take the F align obtained in step 3) as the new F ref , and recursively perform the above steps to obtain the final feature for the subsequent input frame feature map.

[0025] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method steps described in any one of claims 1-5.

[0026] The beneficial effects of the present invention are:

[0027] 1. The panoramic video feature alignment method with multi-frame adaptive fusion of the present invention provides a larger and more flexible receptive field for feature data in feature alignment, and at the same time effectively fuses multi-frame image features, effectively coping with the uneven distribution of the panoramic video picture and the flexible and changeable differences between multi-frame images, and realizing the alignment of the input frame features of the panoramic video to the reference frame. Description of the Drawings

[0028] Figure 1 It is a schematic flowchart of the panoramic video feature alignment method with multi-frame adaptive fusion of the present invention.

[0029] Figure 2 It is a schematic diagram of the calculation of the deformable convolution in step 3 of the present invention. Detailed Embodiment

[0030] The present invention will be further described in detail below through specific embodiments. The following embodiments are only descriptive and not restrictive, and the protection scope of the present invention cannot be limited thereby.

[0031] A panoramic video feature alignment method with multi-frame adaptive fusion, and the steps of the method are as follows Figure 1 shown:

[0032] Step 1): Concatenate the features of the reference frame and the input frame; the specific process of concatenating the features of the reference frame and the input frame is as follows:

[0033] Obtain the reference frame feature map F ref and its input frame feature map F nbor , it is required that the sizes of the reference frame and the input frame feature maps are exactly the same, and the feature maps of the reference frame and the input frame are concatenated along the channel direction into a new feature map:

[0034] Obtain the feature maps of the input frame and the reference frame, and extract features Figure 1 Generally, a deep neural network composed of convolutions is used. The sizes of the input frame and the reference frame feature maps must be exactly the same, the channel depth is 128 or 256, and the width and height are both 256. The input frame and the reference frame are concatenated into a fused feature map according to the channel depth dimension, which is used as the input for subsequent convolutions and the next round of loops. The size W×H×C of the fused feature map is 256×256×256 or 256×256×128.

[0035] Step 2): The concatenated feature map passes through a standard convolution, an activation function, and a dilated convolution to obtain the offset matrix of the input frame feature relative to the reference frame feature in space. The specific process is as follows:

[0036] Step 2.1 The concatenated feature map passes through a standard convolution, and the convolution kernel size is 3×3.

[0037] Step 2.2 The feature map passes through the activation function, the activation function is the Relu activation function, and the activation function expression is as follows:

[0038]

[0039] Step 2.3: The feature map undergoes dilated convolution, and the convolution kernel matrix is ​​K. m×n , the element at position (i, j) in the output offset matrix O is calculated as follows:

[0040]

[0041] Where a is the dilation rate (cavity rate). The larger a is, the larger the receptive field of the original video frame obtained by the output feature is. The channel depth of the offset matrix is ​​2.

[0042] The concatenated feature map is further processed by standard convolution, nonlinear activation function, and dilated convolution. The dilation rate of each cycle is set differently. The larger the dilation rate of dilated convolution is, the greater the perception of the feature map is, but the corresponding features will be rougher. Therefore, the dilation rate of dilated convolution is set larger for later cycles.

[0043] Step 3): According to the offset provided by the offset matrix, the reference frame feature map is deformed convolved to obtain the preliminarily aligned feature map. The specific process is as follows:

[0044] Step 3.1 rounds each element in the offset matrix to get the matrix Offset, that is:

[0045] Offseti,j=Round(Oi,j)

[0046] Step 3.2: Apply the matrix Offset to adjust the standard convolution kernel matrix A. If Offseti,j is [0,0], the position of the convolution kernel corresponding to the position element remains unchanged during convolution; if Offseti,j is [0,1], the convolution kernel corresponding to the position element moves up by 1 pixel during convolution; if Offseti,j is [0,-1], the convolution kernel corresponding to the position element moves down by 1 pixel during convolution; if Offseti,j is [1,0], the convolution kernel corresponding to the position element moves right by 1 pixel during convolution; if Offseti,j is [1,1], the convolution kernel corresponding to the position element moves up by 1 pixel during convolution; If Offseti,j is [1,-1], the convolution kernel moves 1 pixel to the right when performing convolution on the corresponding element; if Offseti,j is [-1,0], the convolution kernel moves 1 pixel to the left when performing convolution on the corresponding element; if Offseti,j is [-1,1], the convolution kernel moves 1 pixel to the left when performing convolution on the corresponding element; if Offseti,j is [-1,-1], the convolution kernel moves 1 pixel to the left when performing convolution on the corresponding element.

[0047] Step 3.3 Apply the adjusted convolution kernel to the reference frame feature map F ref for convolution to obtain the preliminarily aligned feature map F align . The element at the (i, j) position in F align is calculated as follows:

[0048] F align (i,j) = ∑ m ∑ n F ref [i + Oi,j,0 - m, j + Oi,j,1 - n]A(m,n).

[0049] The schematic diagram of the deformable convolution calculation in Step 3) is as Figure 2 shown. The feature map in Step (2) is used as the input of the deformable convolution. The offset matrix of the deformable convolution is obtained by a parallel standard convolution layer, and the output feature map of the deformable convolution is the input of one round of loop.

[0050] Step 4): Recursively perform Step 1), Step 2), and Step 3) several times to obtain the finally aligned reference frame feature map. The specific process is as follows: Take the F align obtained in Step 3) as the new F ref , and recursively perform the above steps to obtain the final feature for the subsequent input frame feature map.

[0051] Concatenate the feature map output by the deformable convolution and the feature map after concatenation in the previous round of loop again, and repeat the above process. After 3 - 5 loops, obtain the feature map of the final input frame aligned according to the reference frame.

[0052] The electronic device of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above - mentioned panoramic video feature alignment method for multi - frame adaptive fusion.

[0053] Although the embodiments and drawings of the present invention are disclosed for illustrative purposes, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the scope of the present invention is not limited to the content disclosed in the embodiments and drawings.

Claims

1. A panoramic video feature alignment method for multi-frame adaptive fusion, characterized by: The following steps are involved: Step 1): Concatenate the features of the reference frame and the input frame; Step 2): The concatenated features are subjected to standard convolution, activation function and dilated convolution to obtain the spatial offset matrix of the input frame features relative to the reference frame features; Step 3): According to the offset provided by the offset matrix, the reference frame feature map is deformed convolved to obtain a preliminarily aligned feature map; Step 4): Recursively perform steps 1), 2), and 3) several times to obtain the final aligned reference frame feature map.

2. The method for panoramic video feature alignment based on multi-frame adaptive fusion according to claim 1, characterized in that: The specific process of splicing the features of the reference frame and the input frame in step 1) is: Get the reference frame feature map F ref The feature map F of its input frame nbor , requiring the size of the reference frame and the input frame feature map to be exactly the same, and concatenating the reference frame and the input frame feature map along the channel direction into a new feature map: F=[F ref ,F nabor ]。 3. The method for panoramic video feature alignment based on multi-frame adaptive fusion according to claim 1, characterized in that: The specific process of obtaining the spatial offset matrix of the input frame features relative to the reference frame features by subjecting the spliced ​​features in step 2) to standard convolution, activation function and dilated convolution is as follows: Step 2.1 The concatenated feature map undergoes standard convolution with a convolution kernel size of 3×3. Step 2.2 The feature map passes through the activation function, the activation function is the Relu activation function, and the activation function expression is as follows: Step 2.3: The feature map undergoes dilated convolution, and the convolution kernel matrix is ​​K. m×n , the element at position (i, j) in the output offset matrix O is calculated as follows: Where a is the expansion rate (hole rate). The larger a is, the larger the receptive field of the original video frame obtained by the output feature is.

4. The method for panoramic video feature alignment based on multi-frame adaptive fusion according to claim 1, characterized in that: In step 3), according to the offset provided by the offset matrix, the specific process of obtaining the feature map after preliminary alignment by deformable convolution of the reference frame feature map is as follows: According to the offset provided by the offset matrix O, the reference frame feature map F ref After deformable convolution, the feature map F is obtained after preliminary alignment align , let the convolution kernel matrix of deformable convolution be A m×n , F align The element at position (i,j) in is calculated as follows: F align (i,j)=∑ m ∑ n F ref [i+O(i,j)-m,j+O(i,j)-n]A(m,n)。 5. The method for panoramic video feature alignment based on multi-frame adaptive fusion according to claim 1, characterized in that: In step 4), steps 1), 2), and 3) are recursively performed several times to obtain the final aligned reference frame feature map. The specific process is: the F obtained in step 3) is align As a new F ref , recursively perform the above steps to obtain the final feature map of the subsequent input frame.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method steps described in any one of claims 1 to 5 are implemented.